Story
arxiv_llm_reliability ยท Oct 8, 2026 ยท paper
arxiv.orgOct 8, 2026
original source linked
In brief
LLM-as-a-Judge is the standard paradigm for NLP evaluation, yet its systemic reliability remains poorly understood despite being widely treated as a deterministic ground truth. We present a comprehensive reliability a...
Feed lens
evaluation