As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fre... Context & related coverage →
CodeScene has published a case study in which coding agents refactored 300,000 lines of C over three weeks for roughly $4,000 in tokens, verified by a frame-by-frame replay harness. The agents built a playbook of code... Context & related coverage →
We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did s... Context & related coverage →
arxiv.org · 2026-09-29 · Ranked: agentic + harness match · research watch · fresh 0.90 · score 2.21
Agent harness bugs exhibit unique characteristics and remain challenging for state-of-the-art software agents to repair. Progress in this area is further hindered by existing benchmarks, which contain only a small and... Context & related coverage →
Vocal · 2026-09-30 · Ranked: harness match · community signal · fresh 0.97 · score 2.12 · Context
Use Jev as a judge for LangSmith evals to evaluate agent traces with faster, cheaper structured feedback across production runs, datasets, and regression tests. Context & related coverage →
arxiv.org · 2026-09-29 · Ranked: evaluation match · research watch · fresh 0.91 · score 2.06
Answering questions about long videos often requires connecting events involving the same objects across hours or days. Chronological descriptions and text-derived entities can leave physical identity unresolved: diff... Context & related coverage →
arxiv.org · 2026-09-29 · Ranked: agent + evaluation match · research watch · fresh 0.91 · score 2.05
Training agents to follow arbitrary instructions is an important goal of multi-task reinforcement learning (RL). Linear temporal logic (LTL) provides a precise and structured formalism for specifying instructions to a... Context & related coverage →
arxiv.org · 2026-09-29 · Ranked: evaluation match · research watch · fresh 0.90 · score 2.02
Photographic evidence is becoming increasingly vulnerable to forms of alteration and fabrication that existing legal and technical workflows are not well equipped to evaluate. Surveillance frames, dashcam stills, and... Context & related coverage →
My comment on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price — Hacker News. I'm a bit late with the pelicans because I was live-blogging the keynote: https://simonwillison.net/2026/Sep/29/openai-devday-... Context & related coverage →
Explore more than 20 announcements from OpenAI DevDay 2026, including GPT-6 Astra, ChatGPT, Codex, APIs, security, and new tools for builders. Context & related coverage →
✓ You're all caught up
Top 12 ranked stories in this snapshot · fresh brief every 2 hours