Can Jev Be a Better Agent Evaluator?
We tested using Jev-as-a-Judge against LLM judges on accuracy, repeatability, latency, and cost to see whether System One models could offer a new approach to agent evaluation. Context & related coverage →
AI news for platform & agent engineers
Ranked signal · finite reading
One shared ranking. Scan what changed, save what matters, and stop when the finish line appears.
Ranked brief · refreshes every 2 hours
We tested using Jev-as-a-Judge against LLM judges on accuracy, repeatability, latency, and cost to see whether System One models could offer a new approach to agent evaluation. Context & related coverage →
Use Jev as a judge for LangSmith evals to evaluate agent traces with faster, cheaper structured feedback across production runs, datasets, and regression tests. Context & related coverage →
In this episode, Sahil Agarwal talks about the critical challenges of identity, authorisation, and security in the age of AI agents. Sahil introduces the DPACT framework (Delegation, Policy, Auditability, Context, and... Context & related coverage →
OpenAI’s Vinoth Govindarajan discusses why production AI agents fail beyond model hallucination. Using real-world case studies like OpenClaw, he explains the key principles of reliable agent harnesses: establishing ex... Context & related coverage →
OpenAI outlines a path to shared global AI standards, calling for coordinated evaluation, reporting, and governance to improve safety. Context & related coverage →
The expanded form of a testimony I prepared for Congress. Context & related coverage →
Every AI factory needs power and cooling that fit its computing architecture. As AI infrastructure expands, power, cooling, water, site and grid constraints are shaping what builders can deploy. Choosing products that... Context & related coverage →
Explore new OpenAI Academy learning paths for employees, developers, leaders, educators, and students to build and demonstrate practical AI skills. Context & related coverage →
OpenAI is working with an independent Advisory Group on Mathematics and Artificial Intelligence to guide the review and communication of emerging AI results. Context & related coverage →
How vLLM reaches 5K throughput and 180 interactivity on Qwen3.8-2.4T with GB300 NVL72 PD serving and how to reproduce results yourself. Context & related coverage →
Using GPT-5.6, V7 turns scattered company files into context agents can use to complete complex, source-linked work. Context & related coverage →