arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MineCEraft:评估语言模型在《我的世界》中作为建造工程师的能力

MineCEraft: Evaluating Language Models as Construction Engineers in the World of Minecraft

Sewoong Lee, Risham Sidhu, Julia Hockenmaier, Yoonhwa Jung

arXiv 2608.28884首次发表:更新:

AI 中文总结

本研究推出开源基准MineCEraft,评估LLMs在《我的世界》建造任务中的表现,通过723条指令、17类任务的实验,分析了LLMs应用于建造工程的失败模式与挑战。

AI 中文摘要

我们推出MineCEraft(Minecraft Construction Engineering Benchmark,发音为mine-see-ee-raft),这是一个易于使用的开源基准,旨在系统评估大型语言模型(LLMs)在《我的世界》建造任务中的可靠性与局限性。该基准包含723条由领域专家手工构建的自然语言指令,具备可程序化验证的评估方式,涵盖17个不同任务类别,为评估LLMs执行真实建造工程任务的能力提供了安全可控的实验环境。我们利用该基准对当前最先进的LLMs进行了深入评估,并开展了详细的错误分析,揭示了LLMs应用于建造工程任务时的关键失败模式与实际挑战。

英文摘要

We introduce MineCEraft (Minecraft Construction Engineering Benchmark, pronounced mine-see-ee-raft), an easy-to-use, open-source benchmark designed to systematically evaluate the reliability and limitations of LLMs for construction tasks in Minecraft. The MineCEraft benchmark comprises 723 domain-expert hand-crafted natural-language instructions with programmatically verifiable evaluation, spanning 17 distinct task categories, providing a safe and controllable experimental environment for assessing LLMs' ability to perform realistic construction engineering tasks. With this benchmark, we conduct an in-depth evaluation of state-of-the-art LLMs and perform a detailed error analysis, revealing key failure modes and practical challenges in applying LLMs to construction engineering tasks.

Journal refEMNLP 2026 Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑