Story
arxiv_cs_cl ยท Jun 3, 2026 ยท paper
Source brief
Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data
arxiv.orgJun 3, 2026
original source linked
In brief
Large language models are increasingly evaluated by other models, raising a natural question: can a model predict how a judge will score its own output? We find that the ability is largely present before any targeted...
Continue reading