Story
arxiv_llm_reliability ยท Aug 27, 2026 ยท paper
arxiv.orgAug 27, 2026
original source linked
In brief
Autoregressive large language models (LLMs) routinely generate factually incorrect outputs with high decoding confidence, limiting their deployment in high-stakes workflows. Existing output-stage uncertainty metrics c...
Continues in
Feed lens
evaluation