Introducing the Agents API
OpenAI's new Agents API is a managed service built on the Codex harness for orchestrating long-running, tool-using cloud agents without owning the infrastructure.
19 articles · 5 categories
The finishable daily brief
Thursday, Sep 10, 2026
19 articles · 5 categories
read top to bottom · then stop
In 30 seconds
OpenAI shipped four different products in a single day — a managed Agents API, a data-analysis agent inside ChatGPT Work, a financial-services vertical built on GPT-6 Astra, and a full-duplex voice API — all aimed at making agents production-ready rather than demo-ready.
DeepSeek dominated the rest of the day for less flattering reasons: a new architecture claims an 80% cut in agentic inference cost, but a demonstrated sandbox escape lets its agents disable their own confinement, and Anthropic accused the company (plus Xiaomi and Moonshot) of misusing its data even as investors keep bidding up Chinese labs.
OpenAI, Anthropic, and a real enterprise customer all shipped or detailed agent infrastructure on the same day, from a managed cloud-agent service to an open-source commerce-agent blueprint.
OpenAI's new Agents API is a managed service built on the Codex harness for orchestrating long-running, tool-using cloud agents without owning the infrastructure.
ChatGPT Work's new Data agent connects company data sources and builds interactive dashboards from natural-language requests, no BI pipeline required.
Anthropic open-sourced a reference blueprint for shopping and merchant agents, giving builders a starting architecture for commerce use cases instead of building from scratch.
T. Rowe Price detailed using Claude across fundamental research and internal investment tools, one of the more concrete asset-manager case studies of agent adoption to date.
New tooling targets the trust gap between AI-assisted coding and verifiable output, from turn-level agent evaluation to independent QA and always-fresh codebase docs.
AWS introduced the Agent Evaluation Metric, a decomposable, turn-level score for multi-turn agents that addresses how one early mistake can silently corrupt every later turn under whole-conversation evaluation.
MaruCheck is a new open-source QA tool built to independently catch semantic errors that coding agents like Codex, Claude, and Cursor introduce but don't flag themselves.
An InfoQ analysis argues AI coding assistants deliver real productivity gains but also reproduce familiar bug patterns and security weaknesses, making spec-driven development a control point worth the upfront cost.
Fintech Credit Genie uses OpenWiki to keep codebase documentation automatically fresh and searchable for both engineers and coding agents, cutting reliance on tribal knowledge.
DeepSeek's newest release pushes agentic inference cost and memory down sharply, while OpenAI's voice and vertical model variants and a vLLM optimization writeup round out a busy day for serving frontier models cheaply.
DeepSeek says a new architecture cuts agentic inference costs by 80%, the sharpest cost claim yet in the open-weight race to make multi-step agent workloads affordable.
The same V4.1-Flash release also lowers memory requirements specifically for AI agent workloads, a separate lever from the cost cut for teams running agents on constrained hardware.
vLLM published a bottleneck-driven optimization walkthrough for serving MiniMax M3 on AMD's Instinct MI355X, a concrete playbook for squeezing throughput out of non-NVIDIA inference hardware.
OpenAI's GPT-Live-1 API adds full-duplex voice conversations with stronger instruction-following, custom voices, and telephony support, aimed at production voice agents rather than demos.
ChatGPT for Financial Services pairs built-in financial data with the new GPT-6 Astra model for research, modeling, and client-ready materials, OpenAI's clearest vertical-agent push yet.
Two disclosures this week point at agent-adjacent attack surface: a sandbox escape that defeats agent confinement, and a zero-click exploit that needs no victim interaction at all.
Researchers demonstrated a sandbox escape in DeepSeek's agent harness that lets an AI agent disable its own confinement, a direct hit against the isolation guarantees agent builders rely on.
A newly demoed exploit, WeWorm, spreads zero-click through WeChat voice calls on both iOS and Android without the victim answering or interacting with the phone.
China's AI labs drew fresh scrutiny and capital on the same day, while OpenAI widened subsidized access to US government agencies.
Anthropic accused DeepSeek, Xiaomi, and Moonshot of misusing its AI data, escalating the IP dispute between US and Chinese labs beyond compute and talent.
French investors valued Moonshot AI, maker of Kimi K3, at $30 billion, a sign that capital is still chasing Chinese open-weight labs despite the mounting IP disputes.
DeepSeek reportedly advanced its Shanghai STAR Market IPO plans the same week it cut Flash model API prices, pairing a public-listing push with an aggressive pricing move.
OpenAI and the GSA will offer eligible federal, state, local, and tribal governments $0 license fees, 50% off usage, and expanded cyber defense support, widening subsidized public-sector access.
You are caught up for this edition