LLM Digest
Subscribe

AI Daily Recap

28 articles · 4 categories

View as JSON
‹

The finishable daily brief

What happened in AI — Oct 2, 2026

Friday, Oct 2, 2026
28 articles · 4 categories

read top to bottom · then stop

In 30 seconds

  • Open SWE's model router cut median cost per coding task by 64% with no measurable quality drop.
  • Docker brings the Sandbox Kit Spec to the CNCF: agent permissions packaged as OCI images.
  • OpenAI DevDay: GPT-6.1 Sol, computer use in the Agents API, cloud Codex environments.
  • Uber Eats cut search latency 50%; AWS shows RL fine-tuning for small search agents.
  • A Chinese open-weight model is under investigation over bioweapon-instruction claims.

Agent cost and safety are being decided in the harness. Open SWE's router cut median coding-task cost 64%, and Docker is taking sandbox permissions to the CNCF as OCI images.

OpenAI's DevDay adds GPT-6.1 Sol and computer use to the Agents API, while local hardware and open-weight safety probes show where self-hosted agents stand.

Agent Harnesses, Routing, and Managed Runtimes 5 items

Harness-level design is where agent cost and safety are now decided: model routing, sandbox permissions, and managed microVMs.

How to Build a Model Router in the Harness

langchain_blogDetails

Open SWE's harness-level model router cut median cost per coding task 64% with no measurable quality drop; the post shows how to build one.

Docker Sandbox Kit Spec: Packaging AI Agent Permissions as OCI Images

infoq_ai_mlDetails

Docker is taking its Sandbox Kit Spec to the CNCF, packaging what an agent may access as an OCI image that travels with the agent.

Frontier Model Releases and Agent Tooling 4 items

OpenAI's DevDay and GPT-6 guidance push reasoning-effort tuning, computer use, and hosted Codex into the production agent path.

Training, Latency, and Production Case Studies 3 items

Teams report measured gains from fine-tuning small agents and rebuilding pipelines rather than adopting bigger models.

Local Inference and Open-Weight Safety Scrutiny 5 items

Local hardware makes open models practical for coding agents, while Chinese open-weight models draw safety and censorship scrutiny.

You are caught up for this edition