Eval Engineering Skill: Build Evals From Repo Context and Traces
LangChain's new skill reads an agent's repo and production traces, proposes evals through user interviews, and outputs runnable Harbor test tasks.
19 articles · 5 categories
The finishable daily brief
Wednesday, Jul 22, 2026
19 articles · 5 categories
read top to bottom · then stop
In 30 seconds
Moonshot's Kimi K3 forced a simultaneous economic and political reckoning today: vLLM shipped production-grade serving support, Microsoft is evaluating swapping ChatGPT and Claude out of Copilot to cut inference costs by $600M, and the White House escalated claims that the model was built on smuggled Nvidia chips and cloned Anthropic IP.
On the practice side, LangChain shipped a skill that generates evals straight from an agent's repo and production traces, and Anthropic published the containment architecture — filesystem, network, and execution limits — it uses to sandbox Claude across Web, Code, and Cowork.
Eval generation and agent architecture both moved from talk to shipped tooling: LangChain now automates eval creation from real traces, and a conference talk argues agents need versioned, composable "virtual tools" instead of ad hoc prompt chains.
LangChain's new skill reads an agent's repo and production traces, proposes evals through user interviews, and outputs runnable Harbor test tasks.
LangChain frames LangGraph's three years as validation that graph-based orchestration, not single-model prompting, is the durable pattern for reliable agents.
Jake Mannix argues agents should move past ad hoc "1970s BASIC" architectures toward an intermediate protocol layer of versioned, encapsulated "virtual tools."
Agent containment moved from research topic to operational priority: Anthropic detailed the sandboxing limits it runs Claude under, a Chinese lab's model was reportedly hacked after copying OpenAI outputs, and analysts flagged cybersecurity as a rising theme across this week's AI coverage.
Anthropic argues agent safety depends on deterministic limits on an agent's filesystem, network, and execution environment, not just model-level alignment.
A report says Zhipu AI's model was compromised after ingesting outputs from misbehaving OpenAI models, an early example of contamination risk from cross-model data reuse.
A roundup notes a cluster of new cybersecurity-focused AI headlines this week, signaling security is becoming a recurring editorial theme rather than a one-off story.
Moonshot's Kimi K3 forced simultaneous engineering, business, and political responses today: vLLM shipped optimized production serving, Microsoft is weighing a swap into Copilot to cut costs, and the White House escalated IP-theft and chip-smuggling accusations against Moonshot.
vLLM previewed production-scale Kimi K3 serving, including KDA-aware prefix caching, fused kernels, optimized MXFP4 MoE, and multimodal support on both NVIDIA and AMD paths.
Microsoft is reportedly evaluating Kimi K3 as a Copilot backend to cut roughly $600M in inference costs versus its current OpenAI and Anthropic model mix.
Interconnects' recap digs into whether Kimi K3 and Qwen 3.8 were distilled from closed frontier models, and what that implies for how fast the open-closed performance gap keeps closing.
A White House official accused Moonshot AI of training Kimi K3 using export-controlled Nvidia chips and technology taken from Anthropic, escalating the dispute beyond generic IP claims.
OpenAI's president publicly called Kimi K3 "pretty good" while declining to confirm whether it was distilled from a closed frontier model, a notably measured reaction from a direct competitor.
A cluster of new tools targets how builders run and pay for coding agents: a self-hosted router to cut per-call cost, two native-Mac control surfaces for running multiple agents at once, and a fresh breakdown of what Copilot's usage-based billing actually buys versus raw API access.
Millwright is a self-hosted, Rust-based LLM router built for cost control and transparency, a response to hosted routers proliferating (and OpenRouter's possible acquisition) without an open alternative.
GitHub breaks down what Copilot's now-listed-API-rate billing buys beyond raw model access: the coding workflow, policy controls, and harness engineering wrapped around the calls.
Rabbitty is a native macOS terminal built specifically to run and manage multiple AI coding agents side by side.
Forkbench is another native-Mac control surface for supervising CLI coding agents, part of a fast-forming category of agent-fleet management tools.
Langy reads a platform's production traces, writes Scenario tests and evaluations for problems it finds, and opens a pull request against the repo on its own.
The frontier labs made three large forward-looking commitments: OpenAI launched an enterprise voice/chat agent platform, Anthropic funded external economic research on AI's labor impact, and Google earmarked compute credits for scientific-discovery research.
OpenAI launched Presence, positioned as a proven enterprise agent platform for deploying trusted voice and chat agents across customer and internal workflows.
Anthropic committed $200 million to fund external research on AI's economic effects, aiming to get independent data ahead of policy debates rather than after them.
Google committed $40M in AI tokens and compute credits to the Genesis Mission, a push to apply frontier models directly to open scientific-discovery problems.
You are caught up for this edition