AWS's Motorway pipeline (Strands Agents + AgentCore) cut wrong agent results from 1 in 8 queries to 1 in 50, and cut incident-detection time from hours to minutes.
LangChain detailed the Harbor-based benchmark it runs across coding, conversation, and retrieval before every Deep Agents change ships.
The White House escalated its claim that Moonshot AI distilled Anthropic's Fable model to build Kimi K3; China called its progress "self-reliance."
Microsoft is still evaluating Kimi K3 for Copilot despite the dispute, and Nvidia's CEO downplayed the competitive threat from Chinese open models.
vLLM shipped an AFD plugin disaggregating attention and FFN for MoE serving across GPU and Ascend NPU backends.
AWS and LangChain each published production eval blueprints today: Motorway's pipeline on Strands Agents and AgentCore cut wrong results from 1 in 8 queries to 1 in 50 and cut incident-detection time from hours to minutes, while LangChain detailed the Harbor-based benchmark it runs across coding, conversation, and retrieval before every Deep Agents change ships.
The White House also escalated its case that Moonshot AI distilled Anthropic's Fable model to build Kimi K3, a claim Beijing rejected as evidence of Chinese "self-reliance" — even as Microsoft confirmed it is still evaluating K3 for Copilot and Nvidia's CEO downplayed the competitive threat.
Agent Engineering: Evals Get Production Numbers 5 items
Two vendors moved eval rigor from talk to measured production outcomes today, and two more benchmarks target the risk side of agent evaluation.
Motorway and AWS built an eval pipeline on Strands Agents and AgentCore that cut wrong results from 1 in 8 queries to 1 in 50 and slashed incident-detection time from hours to minutes.
A production A2A/MCP multi-agent architecture inside a 5G core keeps detection rules aligned with a threat landscape that evolves faster than analysts can write rules by hand.
Agent Tooling: New Harnesses and Frameworks Ship 5 items
Independent builders keep shipping agent harnesses and dev tooling faster than any platform standard has emerged to absorb them.
July's roundup covers the NemoClaw Deep Agents blueprint, a free LangSmith Sandboxes trial, a Fleet Slack integration, and reinforcement-learned-memory support in Deep Agents.
An open-source TypeScript SDK collects privacy-preserving browser trust signals — automation and virtual-camera detection, liveness — without making the block/allow call itself.
Serving Infrastructure and Narrower Model Bets 4 items
Labs and infra vendors each shipped a narrower, more specialized bet today: MoE-optimized serving, dedicated defense compute, a compact model beating a far larger open-weights rival, and consumer health-data integration.
Jensen Huang commissioned a DGX GB300 system at the Naval Postgraduate School, bringing a full-scale AI research platform online for defense-related work.
Poolside's co-CEO describes the small-team model factory behind Laguna S, a 118B-parameter MoE the company says beats a roughly 1T-parameter open-weights rival.
Eligible US users can now connect medical records and Apple Health to ChatGPT for more personalized health insights.
Kimi K3: The Theft Claim Hardens While Business Keeps Moving 5 items
Washington's case that Moonshot AI distilled Kimi K3 from Anthropic's Fable model escalated today, but neither Beijing's rebuttal nor the dispute itself has slowed enterprise interest in the model.
Beijing rejected the theft allegations, framing Kimi K3 and other Chinese models as products of domestic self-reliance rather than copied US technology.
Despite the ongoing dispute, Microsoft is reportedly still evaluating Kimi K3 for use in Copilot, a sign enterprises aren't waiting for the allegations to resolve.
Separately from the origin dispute, a Moonshot model reportedly completed a chip-design task in two days, a concrete capability claim against incumbent EDA tools.