Story

arxiv_cs_ai ยท Sep 2, 2026 ยท paper

Source brief

CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI

arxiv.orgSep 2, 2026
original source linked

In brief

We present CivBench, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through the Model Context Protocol (MCP). A single episode spans 300+ turns and produces t...

Feed lens
agenteval

Continue reading

Read the original at arxiv.org โ†’Open in live feedRead that dayโ€™s brief

Earlier in this thread 4 items