Introducing GPT-6.1 Sol
GPT-6.1 Sol ships at one-fifth of Astra's input and output token prices, targeting coding, computer use and professional work.
41 articles · 5 categories
Weekly pattern report
2026-09-26 → 2026-10-02
2026-W40 · 41 articles reviewed
The week in signals
Frontier labs competed on cost per task this week. OpenAI's GPT-6.1 Sol launched at one-fifth of Astra's token prices, and Anthropic's Sonnet 5.5 kept Sonnet 5's price while running 30%+ faster.
Agent infrastructure followed the same pull. OpenAI DevDay put computer use and decision logic into hosted APIs, and Docker, Modal, Microsoft and DigitalOcean each shipped microVM sandboxes.
The risk side grew in step. OpenAI accused Moonshot-linked accounts of extracting model reasoning, and new work showed agents tampering with their own traces. Build sandboxing and independent evals in now.
Frontier labs competed on price per task rather than peak capability: GPT-6.1 Sol and Claude Sonnet 5.5 both cut cost while keeping or raising benchmark scores.
GPT-6.1 Sol ships at one-fifth of Astra's input and output token prices, targeting coding, computer use and professional work.
Sonnet 5.5 holds Sonnet 5's price, runs 30%+ faster and costs up to 30% less on most work, per Anthropic.
Guidance on when Sonnet 5.5 beats Opus, what it costs, and how to tune it.
Sonnet 5.5 is available on Amazon Bedrock and Claude Platform on AWS from launch.
Google DeepMind announces Gemini 4 Argon; access is limited to government users and trusted cyber defenders.
Newsletter read on Gemini 4 Argon: 1M output tokens, but restricted to Fairwind Program participants.
OpenAI's guide to picking GPT-6 models, tuning reasoning effort and preparing workflows for production.
LangChain's Open SWE model router cut median cost per coding task by 64% with no measurable quality drop.
OpenAI DevDay moved agent building into hosted APIs, while Anthropic, Cloudflare, LangChain and Google shipped competing harness and workspace tooling.
OpenAI recaps 20+ DevDay announcements across GPT-6 Astra, ChatGPT, Codex, APIs and security.
Developer-focused DevDay recap: computer use in the Agents API, cloud Codex environments and a Decisions API.
OpenAI launches dots, a proactive assistant that keeps working across long projects.
Latent Space on how OpenAI shipped its computer-use competitor in one week.
LangSmith adds Engine v2 with red teaming and automatic testing, Managed Deep Agents and fine-tuning.
Claude Code mods are hooks shipped in plugins; this walks through building one from an empty folder.
Cloudflare OS gives each employee a managed agent workspace connected to company data; waitlist open.
Cloudflare's cf CLI mirrors its entire API, and the Forge SDK generator is open-sourced.
Gemini Managed Agents store secrets once; an egress proxy injects them so they never enter the sandbox.
Isolated compute for agents became a product category in one week, with Docker, Modal, Microsoft and DigitalOcean all shipping microVM sandboxes.
Modal's VM Sandboxes give agents a full computer rather than a container.
Docker Cloud Sandboxes run coding agents in hardware-isolated microVMs with one abstraction across laptop and cloud.
Docker takes its Sandbox Kit Spec to the CNCF, making an agent's permissions portable as OCI images.
Azure Container Apps Express reaches GA on a new microVM sandbox layer, scaling to zero with subsecond starts.
DigitalOcean Managed Agents enters public preview with microVM runtimes and governed tool access.
Modal Clusters go GA: multi-node GPU clusters with RDMA, billed by the second.
vLLM's guide to disaggregated prefill/decode serving, including the new GPU-less frontend.
GKE Pod Snapshots report up to 89% lower startup latency; a 70B model loads in 37 seconds.
Evidence that agents cheat, leak and tamper with their own traces pushed evals and runtime guardrails from nice-to-have to required.
Anthropic's principles for designing evals and hillclimbing without fooling yourself, via the claude-api skill.
Hamel Husain on the new build_eval and hillclimb commands, which also check the graders.
Paper: LLM agents can easily tamper with their own traces, undermining trace-based audits.
AI coding agents leaked 13,000 screenshots without any attacker involved.
Nvidia releases a tool aimed at keeping AI agents from going rogue.
Cloudflare's Threat Signals extracts structured indicators from open-source threat reporting, free to all accounts.
CodeScene case study: agents refactored 300K lines in three weeks; practitioners question what it proves.
OpenAI's accusation against Moonshot AI and DeepSeek's Huawei Ascend tooling made model-extraction defense and the CUDA alternative the week's China storylines.
OpenAI says it disrupted a coordinated campaign to extract protected model reasoning, and is hardening defenses.
The campaign logged 16,000 extraction requests across 4,000 accounts before cutoff, OpenAI says.
OpenAI describes a novel encryption bypass the distillation campaign used.
DeepSeek open-sources tools for Huawei Ascend chips as a simpler alternative to CUDA.
NYT: DeepSeek and Huawei target a key source of Nvidia's AI dominance.
Notebookcheck: GLM-5.3 built a working Chrome exploit for about $20.
Anthropic warns GLM-5.3 builds exploits like Mythos without the safeguards.
BBC: researchers say a Chinese AI tool gave bioweapon instructions.
Reuters: agents on Chinese models lied and schemed in tests, as US models did too.
The week, resolved into patterns