How to Build a Model Router in the Harness
Open SWE's harness-level model router cut median cost per coding task 64% with no measurable quality drop; the post shows how to build one.
28 articles · 4 categories
The finishable daily brief
Friday, Oct 2, 2026
28 articles · 4 categories
read top to bottom · then stop
In 30 seconds
Agent cost and safety are being decided in the harness. Open SWE's router cut median coding-task cost 64%, and Docker is taking sandbox permissions to the CNCF as OCI images.
OpenAI's DevDay adds GPT-6.1 Sol and computer use to the Agents API, while local hardware and open-weight safety probes show where self-hosted agents stand.
Harness-level design is where agent cost and safety are now decided: model routing, sandbox permissions, and managed microVMs.
Open SWE's harness-level model router cut median cost per coding task 64% with no measurable quality drop; the post shows how to build one.
Docker is taking its Sandbox Kit Spec to the CNCF, packaging what an agent may access as an OCI image that travels with the agent.
DigitalOcean Managed Agents entered public preview, running agents in isolated microVMs on managed infrastructure.
Pi 1.0 ships as a stable, TypeScript minimalist harness, alongside Pi Durable for long-running work.
A Claude Code UI mod that explains each step, aimed at the opaque shell commands models now run in deep-work mode.
OpenAI's DevDay and GPT-6 guidance push reasoning-effort tuning, computer use, and hosted Codex into the production agent path.
DevDay 2026 brought GPT-6.1 Sol, computer use in the Agents API, and cloud-based Codex environments.
OpenAI's GPT-6 guide covers choosing a model, tuning reasoning effort, and coordinating tools and skills for production workflows.
Chatham Financial used Codex and GPT-5.6 to cut trade validation from 30 minutes to under 4.
Databricks says more than 1 million Genie Agents were created in 2026 and offers guidance on picking first use cases.
Teams report measured gains from fine-tuning small agents and rebuilding pipelines rather than adopting bigger models.
AWS shows multi-turn RL on SageMaker AI teaching a small search agent your tools, targeting frontier-model reliability at lower latency and cost.
Uber Eats rebuilt its search pipeline and reports a 50% cut in end-to-end latency, with new Above-the-Fold measurement.
Anthropic committed $100M to train 10,000 Frontier Deployed Engineers by the end of 2027.
Local hardware makes open models practical for coding agents, while Chinese open-weight models draw safety and censorship scrutiny.
NVIDIA's DGX Spark 64GB targets running increasingly capable open models locally for agent development.
HotHardware's review has a GB10 mini PC running a local Qwen coding agent.
A Chinese AI model is under investigation after a researcher said it gave instructions for bioweapons and assassinations.
Researchers found some topics are off limits inside popular free Chinese-made AI models.
DeepSeek and Huawei are reported to be teaming up to challenge Nvidia's CUDA dominance.
You are caught up for this edition