Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
Opening argument: static correctness checks don't capture real diagnostic reasoning.
3 items · 3 sources · 3 days
Operational story trace
Follow in this browser to see new updates on your Live feed.
Latest change
Databricks ran the first live "Grounded Reasoning Cup," testing AI agents on grounded tasks in real time rather than against a fixed offline benchmark.
A clinical-diagnosis benchmark argued reasoning evals need multi-turn, progressively-disclosed cases instead of single-shot correctness checks. Two weeks later, VAKRA extended that critique to agents, scoring API calls and retrieval together instead of in isolation.
Arc
Opening argument: static correctness checks don't capture real diagnostic reasoning.
Extends the critique to agents, introducing a benchmark that scores API and retrieval reasoning jointly.
First live, grounded evaluation event puts the critique into practice outside offline benchmarks.
Opening argument: static correctness checks don't capture real diagnostic reasoning.
Extends the critique to agents, introducing a benchmark that scores API and retrieval reasoning jointly.
First live, grounded evaluation event puts the critique into practice outside offline benchmarks.
What to watch — open questions
Storylines are threaded mechanically from the feed: stories that share a distinctive anchor across multiple days and sources. Each item links to its original source. The evidence trace, current state, and open questions are written by the editor routine and refreshed whenever a new beat lands.