MirrorCraft:Minecraft中隐藏规则变化下的配对评估
MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft
浏览论文内容
中文总结 AI 辅助
MirrorCraft是用于Minecraft隐藏规则变化下评估智能体的配对基准,实验发现规则集对隐藏规则变化的影响差异显著,未提供规则描述时ReAct表现最优,提供规则可适度提升任务表现。
中文摘要 AI 辅助
随着大语言模型(LLM)的兴起,基于LLM的智能体在Minecraft中的表现成为有趣的研究课题。遗憾的是,现有大多数基准测试在固定游戏机制下评估这些智能体,而在这类设置中的高性能并不能表明智能体在熟悉的配方、掉落物及其他规则发生变化时仍能持续进步。本文提出MirrorCraft,一种用于在Minecraft隐藏规则变化下评估智能体的配对基准。每个Mirror世界是其配对Vanilla世界的副本,对应数据包会修改选定的服务器端规则;在每一组Vanilla-Mirror配对中,地形、生成点、资源分布、目标、界面及动作预算保持一致。MirrorCraft包含5种受控生物群系、6套规则集、3种进度目标、2类模型家族,以及在共享Mineflayer接口下的6种智能体配置。我们通过确定性进度里程碑和成功率评估任务进展,并使用规则干预效应(Rule Intervention Effect, RIE)衡量配对Vanilla与Mirror世界间的性能变化。实验表明,不同规则集对隐藏规则变化的影响差异显著;在未提供规则描述的评估配置中,ReAct取得最高的汇总Mirror分数,而提供精确规则可使所有3种目标的平均进度和完成率获得适度提升。MirrorCraft将Minecraft评估扩展至固定机制之外,为研究智能体在当前世界规则与熟悉规则不同时如何利用游戏结果提供了受控环境。
英文摘要
With the prosperity of the large language models (LLMs), it has become an interesting topic: how do LLM-based agents work in Minecraft? Unfortunately, most existing benchmarks evaluate them under fixed game mechanics. High performance in these settings does not show whether an agent can continue making progress when familiar recipes, drops, and other rules change. In this paper, we introduce MirrorCraft, a paired benchmark for evaluating agents under hidden rule changes in Minecraft. Each Mirror world is a copy of its paired Vanilla world, with selected server-side rules modified by the corresponding datapack. Terrain, spawn, resource placement, objective, interface, and action budget remain matched within every Vanilla-Mirror pair. MirrorCraft includes five controlled biomes, six rule suites, three progression objectives, two model families, and six agent configurations under a shared Mineflayer interface. We evaluate task progress with deterministic advancement milestones and success rate and use the Rule Intervention Effect (RIE) to measure the performance change between matched Vanilla and Mirror worlds. The experiments show that hidden rule changes have strongly different effects across suites. Among the configurations evaluated without rule descriptions, ReAct achieves the highest pooled Mirror score. Providing the exact rules yields modest gains in average progress and completion across all three objectives. MirrorCraft extends Minecraft evaluation beyond fixed mechanics and provides a controlled setting for studying how agents use gameplay outcomes when the rules of the current world differ from familiar ones.