Story
arxiv_llm_reliability ยท Sep 28, 2026 ยท paper
Source brief
Does Model Uncertainty Track Human Ambiguity? Evidence from Multi-Annotator Vision Benchmarks
arxiv.orgSep 28, 2026
original source linked
In brief
Human-model alignment is critical for trustworthy AI-assisted decision-making systems. Yet, most work evaluates model predictions against single ground-truth labels, overlooking that humans themselves often disagree o...
Feed lens
eval