We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did s... Context & related coverage →
github.com · 2026-09-29 · Ranked: agent + eval match · community signal · fresh 0.99 · score 2.37 · Context
Use Jev as a judge for LangSmith evals to evaluate agent traces with faster, cheaper structured feedback across production runs, datasets, and regression tests. Context & related coverage →
I'm at OpenAI DevDay today, in Fort Mason, San Francisco. Same as last year I'll be live blogging the keynote and some other notes during the day. OpenAI gave me a free ticket and a seat in the "creator" area for the... Context & related coverage →
BBC · 2026-09-29 · Ranked: community signal · fresh 0.99 · score 2.00 · Context
Explore more than 20 announcements from OpenAI DevDay 2026, including GPT-6 Astra, ChatGPT, Codex, APIs, security, and new tools for builders. Context & related coverage →
Meet GPT-6.1 Sol: near-Astra intelligence for coding, computer use, and professional work at one-fifth of Astra’s standard API input and output token prices. Context & related coverage →
latent.space · 2026-09-29 · Ranked: claude code match · practitioner analysis · fresh 0.81 · score 1.56