Story
arxiv_llm_reliability ยท Sep 3, 2026 ยท paper
Source brief
KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents
arxiv.orgSep 3, 2026
original source linked
In brief
As LLMs increasingly act through tools, they must reconcile user instructions, parametric knowledge, and dynamic environmental observations before taking actions. We introduce KC-Bench, a controlled multi-turn benchma...
Feed lens
agentevaluation