Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
First paper: scores multi-turn diagnostic reasoning as clinical info is disclosed progressively, not all at once.
3 items · 2 sources · 2 days
Operational story trace
Follow in this browser to see new updates on your Live feed.
Latest change
An Aug 11 benchmark shows multimodal LLMs' cognitive reliability drops sharply in complex urban scenes under adverse conditions, even though the same models look strong on benign inputs.
Three arXiv papers from late July into August 2026 probe where multimodal LLM reasoning breaks: clinical diagnostic QA that unfolds information progressively rather than all at once, question answering over irregular clinical time series, and now trustworthiness in complex, adverse-condition urban scenes.
Arc
First paper: scores multi-turn diagnostic reasoning as clinical info is disclosed progressively, not all at once.
Same-day companion paper: a cost-effective framework for QA over irregular clinical time series.
Two-week-later followup: a benchmark showing multimodal reliability drops sharply in complex, adverse-condition urban scenes.
First paper: scores multi-turn diagnostic reasoning as clinical info is disclosed progressively, not all at once.
Same-day companion paper: a cost-effective framework for QA over irregular clinical time series.
Two-week-later followup: a benchmark showing multimodal reliability drops sharply in complex, adverse-condition urban scenes.
What to watch — open questions
Storylines are threaded mechanically from the feed: stories that share a distinctive anchor across multiple days and sources. Each item links to its original source. The evidence trace, current state, and open questions are written by the editor routine and refreshed whenever a new beat lands.