AI 中文总结
本研究推出开源基准MineCEraft,评估LLMs在《我的世界》建造任务中的表现,通过723条指令、17类任务的实验,分析了LLMs应用于建造工程的失败模式与挑战。
AI 中文摘要
我们推出MineCEraft(Minecraft Construction Engineering Benchmark,发音为mine-see-ee-raft),这是一个易于使用的开源基准,旨在系统评估大型语言模型(LLMs)在《我的世界》建造任务中的可靠性与局限性。该基准包含723条由领域专家手工构建的自然语言指令,具备可程序化验证的评估方式,涵盖17个不同任务类别,为评估LLMs执行真实建造工程任务的能力提供了安全可控的实验环境。我们利用该基准对当前最先进的LLMs进行了深入评估,并开展了详细的错误分析,揭示了LLMs应用于建造工程任务时的关键失败模式与实际挑战。
英文摘要
We introduce MineCEraft (Minecraft Construction Engineering Benchmark, pronounced mine-see-ee-raft), an easy-to-use, open-source benchmark designed to systematically evaluate the reliability and limitations of LLMs for construction tasks in Minecraft. The MineCEraft benchmark comprises 723 domain-expert hand-crafted natural-language instructions with programmatically verifiable evaluation, spanning 17 distinct task categories, providing a safe and controllable experimental environment for assessing LLMs' ability to perform realistic construction engineering tasks. With this benchmark, we conduct an in-depth evaluation of state-of-the-art LLMs and perform a detailed error analysis, revealing key failure modes and practical challenges in applying LLMs to construction engineering tasks.
Journal refEMNLP 2026 Findings