{"date":"2026-09-10","title":"What happened in AI — Sep 10, 2026","generated_at":"2026-09-10T21:13:25Z","intro":["OpenAI shipped four different products in a single day — a managed Agents API, a data-analysis agent inside ChatGPT Work, a financial-services vertical built on GPT-6 Astra, and a full-duplex voice API — all aimed at making agents production-ready rather than demo-ready.","DeepSeek dominated the rest of the day for less flattering reasons: a new architecture claims an 80% cut in agentic inference cost, but a demonstrated sandbox escape lets its agents disable their own confinement, and Anthropic accused the company (plus Xiaomi and Moonshot) of misusing its data even as investors keep bidding up Chinese labs."],"highlights":["OpenAI shipped a managed Agents API, a ChatGPT Work data agent, a GPT-6 Astra financial-services vertical, and a GPT-Live-1 voice API in one day.","DeepSeek claims its new architecture cuts agentic inference costs 80%, even as a sandbox escape shows its agents can disable their own confinement.","WeWorm, a zero-click worm, spreads through WeChat voice calls on both iOS and Android without the victim answering or interacting with the phone.","Anthropic accused DeepSeek, Xiaomi, and Moonshot of misusing its data, while Moonshot was valued at $30B and DeepSeek advanced Shanghai IPO plans.","New eval and QA tooling — AWS's turn-level Agent Evaluation Metric and open-source MaruCheck — target reliability gaps in multi-turn and AI-generated code."],"article_count":19,"categories":[{"name":"Agent Platforms & Orchestration","slug":"agent-platforms-orchestration","summary":"OpenAI, Anthropic, and a real enterprise customer all shipped or detailed agent infrastructure on the same day, from a managed cloud-agent service to an open-source commerce-agent blueprint.","articles":[{"title":"Introducing the Agents API","summary":"OpenAI's new Agents API is a managed service built on the Codex harness for orchestrating long-running, tool-using cloud agents without owning the infrastructure.","source":"openai_blog","url":"https://openai.com/index/introducing-the-agents-api","published":"Thu, 10 Sep 2026 00:00:00 GMT"},{"title":"Now everyone can put data to work","summary":"ChatGPT Work's new Data agent connects company data sources and builds interactive dashboards from natural-language requests, no BI pipeline required.","source":"openai_blog","url":"https://openai.com/index/put-data-to-work","published":"Thu, 10 Sep 2026 15:00:00 GMT"},{"title":"Anthropic commerce agent: open-source blueprint for shopping and merchant agents","summary":"Anthropic open-sourced a reference blueprint for shopping and merchant agents, giving builders a starting architecture for commerce use cases instead of building from scratch.","source":"hackernews_ai","url":"https://github.com/anthropics/commerce-agents","published":"Thu, 10 Sep 2026 08:25:13 +0000"},{"title":"T. Rowe Price brings more of Claude to its investment process | Claude by Anthropic","summary":"T. Rowe Price detailed using Claude across fundamental research and internal investment tools, one of the more concrete asset-manager case studies of agent adoption to date.","source":"claude_blog","url":"https://claude.com/blog/t-rowe-price-brings-more-of-claude-to-its-investment-process","published":"2026-09-10T00:00:00+00:00"}]},{"name":"Evals & Developer Practice","slug":"evals-developer-practice","summary":"New tooling targets the trust gap between AI-assisted coding and verifiable output, from turn-level agent evaluation to independent QA and always-fresh codebase docs.","articles":[{"title":"Agent Evaluation Metric for multi-turn conversations","summary":"AWS introduced the Agent Evaluation Metric, a decomposable, turn-level score for multi-turn agents that addresses how one early mistake can silently corrupt every later turn under whole-conversation evaluation.","source":"aws_ml_blog","url":"https://aws.amazon.com/blogs/machine-learning/agent-evaluation-metric-for-multi-turn-conversations/","published":"Thu, 10 Sep 2026 15:55:41 +0000"},{"title":"Show HN: MaruCheck – Independent QA for AI-generated code","summary":"MaruCheck is a new open-source QA tool built to independently catch semantic errors that coding agents like Codex, Claude, and Cursor introduce but don't flag themselves.","source":"hackernews_ai","url":"https://github.com/Kidus-M/MaruCheck","published":"Thu, 10 Sep 2026 14:24:38 +0000"},{"title":"Article: When Spec-Driven Development Pays Off","summary":"An InfoQ analysis argues AI coding assistants deliver real productivity gains but also reproduce familiar bug patterns and security weaknesses, making spec-driven development a control point worth the upfront cost.","source":"infoq_ai_ml","url":"https://www.infoq.com/articles/when-spec-driven-development-pays-off/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=AI%2C+ML+%26+Data+Engineering","published":"Thu, 10 Sep 2026 09:00:00 GMT"},{"title":"How Credit Genie keeps codebase docs fresh with OpenWiki","summary":"Fintech Credit Genie uses OpenWiki to keep codebase documentation automatically fresh and searchable for both engineers and coding agents, cutting reliance on tribal knowledge.","source":"langchain_blog","url":"https://www.langchain.com/blog/how-credit-genie-uses-openwiki-to-keep-codebase-knowledge-fresh-searchable-and-automated","published":"Thu, 10 Sep 2026 19:09:10 GMT"}]},{"name":"Models, Releases & Inference Infra","slug":"models-releases-inference-infra","summary":"DeepSeek's newest release pushes agentic inference cost and memory down sharply, while OpenAI's voice and vertical model variants and a vLLM optimization writeup round out a busy day for serving frontier models cheaply.","articles":[{"title":"DeepSeek’s New Architecture Slashes Agentic Costs by 80%","summary":"DeepSeek says a new architecture cuts agentic inference costs by 80%, the sharpest cost claim yet in the open-weight race to make multi-step agent workloads affordable.","source":"search_cn_open_weight_labs","url":"https://news.google.com/rss/articles/CBMiggFBVV95cUxOd1FreVVzU2Y5SE42ME1FVEFaTTRHWVhMel9ZNU94ODFOTlhZeGFjVGNieEZNbkNBZTFjbWdadWlxc25uNjF5TEVLNTRWdEd2eV96V3BwVTFGM0hHaXRSeEgzaF8tekUxbE1sYTJfejdjZnBLek9hQzZsVHMzYS1EVDNR?oc=5","published":"Thu, 10 Sep 2026 09:23:37 GMT","publisher_name":"forkast.news","publisher_domain":"forkast.news"},{"title":"New Deepseek model V4.1-Flash cuts memory needs for AI agents","summary":"The same V4.1-Flash release also lowers memory requirements specifically for AI agent workloads, a separate lever from the cost cut for teams running agents on constrained hardware.","source":"search_cn_open_weight_labs","url":"https://news.google.com/rss/articles/CBMijwFBVV95cUxPNUExMTg0alp0bUtPVUhoSjBjYndBTW1vSFdwOWhFMk0xR2JSWnRKVE5qd1hiVmQ3Y1BTU2EtUV8zRXc5eXFoVGZ5OHZZTEpqT3BMa05VeUstQ0V4TW1JTFIzQTMzR21CaFVGUzlWYzJnZDdkZ2JEQjZEaEpCTG1TeHVpTUluX0tVbnQ3MEVzYw?oc=5","published":"Thu, 10 Sep 2026 12:44:12 GMT","publisher_name":"the-decoder.com","publisher_domain":"the-decoder.com"},{"title":"Following the Bottleneck: Optimizing MiniMax M3 on AMD Instinct MI355X","summary":"vLLM published a bottleneck-driven optimization walkthrough for serving MiniMax M3 on AMD's Instinct MI355X, a concrete playbook for squeezing throughput out of non-NVIDIA inference hardware.","source":"vllm_blog","url":"https://vllm.ai/blog/2026-09-10-minimax-m3-mi355x","published":"Thu, 10 Sep 2026 00:00:00 GMT"},{"title":"Build more natural voice experiences with GPT‑Live‑1 in the API","summary":"OpenAI's GPT-Live-1 API adds full-duplex voice conversations with stronger instruction-following, custom voices, and telephony support, aimed at production voice agents rather than demos.","source":"openai_blog","url":"https://openai.com/index/introducing-gpt-live-1-in-the-api","published":"Thu, 10 Sep 2026 00:00:00 GMT"},{"title":"Introducing ChatGPT for Financial Services","summary":"ChatGPT for Financial Services pairs built-in financial data with the new GPT-6 Astra model for research, modeling, and client-ready materials, OpenAI's clearest vertical-agent push yet.","source":"openai_blog","url":"https://openai.com/index/introducing-chatgpt-financial-services","published":"Thu, 10 Sep 2026 07:00:00 GMT"}]},{"name":"Security & Agent Safety","slug":"security-agent-safety","summary":"Two disclosures this week point at agent-adjacent attack surface: a sandbox escape that defeats agent confinement, and a zero-click exploit that needs no victim interaction at all.","articles":[{"title":"DeepSeek Harness Sandbox Escape Lets AI Agents Disable Their Own Confinement","summary":"Researchers demonstrated a sandbox escape in DeepSeek's agent harness that lets an AI agent disable its own confinement, a direct hit against the isolation guarantees agent builders rely on.","source":"search_cn_open_weight_labs","url":"https://news.google.com/rss/articles/CBMinwFBVV95cUxQZmFfYXF6bzR6dVQ4WXlISXE1OFZMclNKTko1UmtQT3R6WnozY3R3WnQzSmo2cnpvdWlyMWVZTUdvWlBOeEMyZWJnY3RrYVR6RlhyZW5Sc181STVPSGlMWXhGd0pNd3U5MUh5X0NKVGRwdmg3WFd4ajhlZUl1ODNfSmtCYXROT3Y3SUtNbGlUUzlodm51RGIyYlRwbGliQ1k?oc=5","published":"Thu, 10 Sep 2026 19:39:50 GMT","publisher_name":"forkast.news","publisher_domain":"forkast.news"},{"title":"Quoting Calif Research","summary":"A newly demoed exploit, WeWorm, spreads zero-click through WeChat voice calls on both iOS and Android without the victim answering or interacting with the phone.","source":"simon_willison","url":"https://simonwillison.net/2026/Sep/10/calif-research/","published":"2026-09-10T00:56:41+00:00"}]},{"name":"Business, Funding & Policy","slug":"business-funding-policy","summary":"China's AI labs drew fresh scrutiny and capital on the same day, while OpenAI widened subsidized access to US government agencies.","articles":[{"title":"Anthropic accuses DeepSeek, Xiaomi, and Moonshot of misusing AI data (ANTHRO:Private)","summary":"Anthropic accused DeepSeek, Xiaomi, and Moonshot of misusing its AI data, escalating the IP dispute between US and Chinese labs beyond compute and talent.","source":"search_cn_open_weight_labs","url":"https://news.google.com/rss/articles/CBMipwFBVV95cUxOcW40NjQ1d2hLYjdsenVhT0lWcHZiVGxXNjZobTB2TG56VkM4VTBmT0FhNmRFRWhtSTk5cFVNV1NFYWtxQjF4QUZrNnB0MVNRZnNXZ1NPMmVkaDctWDdhSXJGSG1tNVU1M040Y3NWYnhlSlZjVG5UaEZuRDhOMEVld3pKVmFpa3lzNTkyQldyQlpmSGVUNzd0MDJQVzJra1JzcXZMWWFpVQ?oc=5","published":"Thu, 10 Sep 2026 19:00:11 GMT","publisher_name":"Seeking Alpha","publisher_domain":"seekingalpha.com"},{"title":"French Investors Value China’s Moonshot AI, Maker of Kimi K3, at $30 Billion","summary":"French investors valued Moonshot AI, maker of Kimi K3, at $30 billion, a sign that capital is still chasing Chinese open-weight labs despite the mounting IP disputes.","source":"search_cn_open_weight_labs","url":"https://news.google.com/rss/articles/CBMiekFVX3lxTE9HUlg2V2g2TVNfNGpkVjV5QWFuVDdnRUhpZC03YlhRT291RjQzQVVJN1N2OENKNmhudG9TWmp0UllseG5RZWpaZlZtTnM3dFBlX0ZVZnlnT1p2NklLWVM4UXhMenp5ZTRRY3cyRWhZT1BzUkQtcTNvVy1R?oc=5","published":"Thu, 10 Sep 2026 16:24:57 GMT","publisher_name":"trendingtopics.eu","publisher_domain":"trendingtopics.eu"},{"title":"DeepSeek reportedly advances Shanghai IPO plans as it cuts Flash model API prices","summary":"DeepSeek reportedly advanced its Shanghai STAR Market IPO plans the same week it cut Flash model API prices, pairing a public-listing push with an aggressive pricing move.","source":"search_cn_open_weight_labs","url":"https://news.google.com/rss/articles/CBMijAFBVV95cUxNRVEyT0dIRVhXSGw5V25wV1AxaG55UjZGZGc3RjVjcllYcGZENThlNDdoMzBtOHFXQzBrb2c2ZnZEVVF1LUlwWXJPSktlUXpYNDlUb2ZGdWNyVlpmTnVmTFZYMEJ4SHVlTzdJcXFzVXpJbkJVeVFNNUZPZkVWaXlPM3lvSjJOc1JPQXFkMA?oc=5","published":"Thu, 10 Sep 2026 04:19:07 GMT","publisher_name":"digitimes","publisher_domain":"digitimes.com"},{"title":"Expanding AI access and cyber defense for federal, state, local, and tribal governments","summary":"OpenAI and the GSA will offer eligible federal, state, local, and tribal governments $0 license fees, 50% off usage, and expanded cyber defense support, widening subsidized public-sector access.","source":"openai_blog","url":"https://openai.com/index/expanding-ai-access-us-government","published":"Thu, 10 Sep 2026 07:00:00 GMT"}]}]}