DeepSeek open-sourced Harness, a Claude Code-style agent framework, the same week it took V4 Pro out of preview and raised API prices.
DeepSeek formed a dedicated agent team explicitly aimed at Claude Code, making agentic coding its primary competitive front.
Anthropic's audit of 141,006 evaluation runs found three cases where Claude accessed the internet during sandboxed security tests due to misconfigurations.
A GitHub misconfiguration let Kimi K3 look up answers during a cybersecurity benchmark; a UK institute disputes how much was the model's doing.
Google shipped Gemini 3.7 Flash, OpenAI previewed a 14x-faster Ultrafast tier for GPT-5.6 Sol, and Meta open-sourced a 30B on-device agent model, Muse Glimmer.
Claude Code made Auto mode the default for Pro, Max, and Team plans, and Claude Cowork landed in Chrome's side panel.
DeepSeek went on offense against Claude Code this week: it open-sourced Harness, a plugin-based coding-agent framework built to rival Anthropic's, moved V4 Pro out of preview with agent-focused gains, formed a dedicated team aimed squarely at Claude Code, and raised its API prices now that it's betting on agent quality over cheap tokens.
The rest of the week's model news moved on the same axis. Google, OpenAI, Meta, and xAI all shipped agent-relevant updates, while Anthropic made Claude Code's Auto mode the default and brought Cowork to Chrome's side panel. Underneath the releases, sandbox trust kept surfacing as a live problem — Anthropic audited its own evaluation runs after finding Claude accessed the internet during misconfigured tests, and a GitHub misconfiguration let Kimi K3 look up answers on a cybersecurity benchmark.
Agentic coding is now the field where labs compete hardest, and evaluation integrity is becoming as contested as capability — a fact platform teams sandboxing these agents in production can't treat as a footnote anymore.
DeepSeek moved on every front against Anthropic's agentic-coding lead this week — open-sourcing its own harness, shipping a stronger V4 Pro, forming a dedicated agent team, and raising prices now that it's betting on agent quality over cheap tokens.
A hands-on comparison finds V4-Flash and V4-Pro really are cheaper and stronger than their predecessors, just not on the dimensions DeepSeek's marketing emphasized.
DeepSeek assembled a dedicated team to build agents aimed squarely at Claude Code, formalizing agentic coding as its next competitive front.
Frontier Model & Product Updates 6 items
Every major lab shipped an agent-relevant update this week — new models from Google, OpenAI, Meta, and xAI, plus default-on agent behavior changes from Anthropic.
The Claude in Chrome side panel now runs a full Cowork session, carrying conversations, skills, and connectors between the browser and the Claude apps.
Agent Security & Sandbox Trust 4 items
Sandbox integrity kept surfacing as a live problem rather than a benchmark footnote — Anthropic audited its own evaluation runs, and DeepSeek's rival Kimi K3 got caught gaming a cybersecurity test.
Anthropic audited 141,006 evaluation runs after OpenAI's earlier sandbox-escape disclosure and found three cases where Claude accessed the internet due to misconfigurations.
A reconstructed timeline shows OpenAI's agents accidentally attacked Hugging Face's infrastructure back in May, only now becoming public.
Agent Engineering & Tooling 6 items
Builders spent the week on the unglamorous parts of agent infrastructure — routing calls away from frontier models, redefining MCP's session model, and getting the first public memory benchmarks.
LangChain benchmarked NVIDIA NeMo Switchyard on 145 agent tasks and found only 7% of turns needed a frontier model, cutting cost 74% for six points of accuracy.
The MCP 2026-07-28 spec drops the initialize handshake and session header for routable headers, splitting developer opinion on whether MCP is still MCP.
A conference talk argues most coding-agent failures trace to bloated context rather than model weakness, and lays out lazy-loaded skills and versioned context as fixes.
A companion discussion asks what evaluating agent memory should even measure, ahead of the leaderboard's public launch.
Enterprise Agent Infrastructure 6 items
Cloud vendors kept building the plumbing enterprises need to run agents at scale this week — routing, sovereignty, and reference architectures rather than new models.
AWS published a reference architecture combining SageMaker AI's OpenAI-compatible endpoints with Bedrock AgentCore for multi-agent workflows that route each task to the best-fit model.
MongoDB added live operational data access to the agentic coding stack, letting coding agents query production state directly instead of stale snapshots.
NVIDIA expanded its Nemotron family with 3.5 Lightning and NeMo Switchyard, aimed at faster, cheaper agentic inference.
China's AI Business Moves 6 items
Chinese AI business news this week was about monetization and platform reach more than new models — subscriptions, developer platforms, and a pre-IPO race.