CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI
We present CivBench, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through the Model Context Protocol (MCP). A single episode spans 300+ turns and produces thousands of tool calls over a large action space, requiring sustained planning, state monitoring, and execution under partial observability. The environment exposes 76 MCP tools and…
We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.