Live feed
AI news for platform & agent engineers
The daily paper for AI engineers.
Ranked brief · refreshes every 2 hours
Top signals · Oct 9–11, 2026 · Updated Oct 11, 04:04 UTC
Android Bench 2 Adds Support for Long-Horizon Tasks, Agentic Evaluation, and Continuous Scoring
Google has released Android Bench 2.0, a major update to its benchmark framework for evaluating AI models and agents on Android development tasks. The update introduces long-horizon tasks (LHTs), agent-based evaluatio... Context & related coverage →
Investigating unintended model actions in our evaluations and internal use
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems. Context & related coverage →
Quoting The New York Times
Anthropic detailed the activity of its A.I. agents in a blog post on Friday, without naming the targeted websites. But two sources with knowledge of the incidents said Anthropic’s A.I. agents had submitted 20 visa app... Context & related coverage →
[Subscriber Exclusive] NYC Subscriber Meetups!
Tomorrow and Tuesday. If you see this you’re in! Context & related coverage →
Cloudflare Traces Turns the Proxy Layer into OpenTelemetry Spans, with New Volume-Based Pricing
Cloudflare has put Traces into open beta, extending automatic tracing from Workers to security rules, transformations, cache, routing and origin handling as OpenTelemetry spans. Traces accept and forward W3C tracepare... Context & related coverage →
How to build great out-of-the-box user experiences with Managed Deep Agents
Managed Deep Agents includes a new API for managing reactions for your distributed agents, and a system to dynamically assign emoji responses with your instrument of choice. Learn more. Context & related coverage →
[AINews] TypeSafe/Jev at >$100M ARR, $7.5B valuation 3 weeks after launch
Impactful scheduling for GPU clusters
How to Build a Model Router in the Harness
How we built a model router into Open SWE's harness that cut median cost per coding task by 64% with no measurable drop in quality, and how to build your own. Context & related coverage →
Asana cuts model costs 76x in browser tests with GPT-6.1 Sol
Using GPT-6 Astra in Codex, Asana made its browser agent 76x cheaper and 5x faster in tests to offer customers more capable models. Context & related coverage →
Sophos cuts threat investigation time by 96% with OpenAI Daybreak
Discover how Sophos uses OpenAI’s Daybreak to cut cyber-threat investigation time by 96% and automate 52% of MDR cases while preserving human oversight. Context & related coverage →