发表机构
University of Oxford; Tony Blair Institute for Global Change; Google DeepMind; UK AI Security Institute; Imperial College London; Downing Street(牛津大学; 托尼·布莱尔全球变化研究所; 谷歌DeepMind; 英国人工智能安全研究所; 帝国理工学院; 唐宁街10号)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究人员提出开源基准CivBench,用于通过MCP评估语言模型智能体在《文明VI》长周期工具环境中的表现,发现智能体存在战略状态监控不足、近期承诺执行率低的问题。
AI 中文摘要
我们提出CivBench,这是一个开源基准测试,用于通过模型上下文协议(MCP)评估语言模型智能体在长周期、工具介导环境中的表现。单个回合包含300多个步骤,在庞大的动作空间中产生数千次工具调用,要求智能体在部分可观测性下进行持续规划、状态监控与执行。该环境提供76个MCP工具,以及将视觉游戏状态转换为结构化文本的叙事层。我们使用CivBench在23次合规运行中表征四个模型系列的智能体行为。该样本为试点性质,并非模型排名:在此规模下,汇总结果无法可靠区分模型。相反,我们引入了两个环境可量化的接口级指标:主动监控率(PMR),用于衡量智能体是否主动查询潜在战略状态;RAG@10,用于衡量结构化规划反思中所述的承诺是否在后续10个步骤内执行。在共享剧本协议下,我们在所有运行中观察到两个一致模式:智能体对可获取但需显式查询的战略相关状态监控不足,尽管剧本指导每20步查询一次胜利进度,但智能体仅每30至75步查询一次,且在20次可检测的失败中,有7次未能在游戏结束前的20步警告窗口内进行查询;智能体还经常无法执行自身规划反思中所述的近期承诺(各模型的RAG@10介于48.2%至65.8%之间)。尽管具备工具访问权限和明确指导,这两种模式仍会出现,我们将其解释为指令下的偏差而非能力缺失。我们在该httpsURL发布了环境、场景、日志、指标及分析流水线。
英文摘要
We present CivBench, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through the Model Context Protocol (MCP). A single episode spans 300+ turns and produces thousands of tool calls over a large action space, requiring sustained planning, state monitoring, and execution under partial observability. The environment exposes 76 MCP tools and a narration layer that converts visual game state into structured text. We use CivBench to characterise agent behaviour across four model families in 23 admissible runs. The sample is a pilot, not a model ranking: aggregate outcomes do not reliably discriminate models at this scale. Instead, we introduce two interface-level metrics that the environment makes measurable: Proactive Monitoring Rate (PMR), capturing whether agents actively query latent strategic state, and RAG@10, capturing whether commitments stated in structured planning reflections are executed within ten subsequent turns. Across runs we observe two consistent patterns under a shared playbook protocol. Agents under-monitor strategically relevant state that is available but requires explicit querying: despite playbook guidance to query victory progress every 20 turns, agents do so only every 30 to 75 turns, and in 7 of 20 detectable defeats they failed to query within the 20 turn warning window before game end. Agents also frequently fail to execute near-term commitments stated in their own planning reflections (RAG@10 between 48.2% and 65.8% across models). Both patterns arise despite tool access and explicit guidance, and we interpret them as deviations under instruction rather than absences of capability. We release the environment, scenarios, logs, metrics, and analysis pipeline at https://github.com/lmwilki/civ6-mcp