arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ArtifactArena:通过模型在物理世界中的构建物来评估模型

ArtifactArena: Evaluating Models by What They Build in the Physical World

Kushagra Tiwary*, David Mayo*, Nikhil Behari, Xiangzhou Sun, Abdulrahman Alabdulkareem, Isaac Galatzer-Levy, Boris Katz, Brian Cheung

arXiv 2610.06511首次发表:更新:

发表机构

MIT CSAIL; MIT Media Lab; NYU Grossman School of Medicine; UCSF(麻省理工学院计算机科学与人工智能实验室; 麻省理工学院媒体实验室; 纽约大学格罗斯曼医学院; 加州大学旧金山分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出ArtifactArena平台,通过物理模拟竞技场中的机器人构建竞赛,评估前沿模型的零样本、验证器引导和开放式物理设计能力,并以Elo排名基准测试。

AI 中文摘要

为了评估前沿能力,我们必须衡量模型不是根据它们说了什么,而是根据它们在接地物理环境中能够设计和构建什么。我们引入了ArtifactArena,一个开放式平台,模型在此面临物理接地硬件-软件协同设计挑战:工程完全功能的机器人以在模拟竞技场中竞争。我们通过三个测试工具评估前沿模型的零样本、验证器引导的改进以及开放式物理设计能力,这些工具基于文本描述、物理模拟器反馈和游戏数据来改进它们的机器人工件。我们通过基于工件之间头对头锦标赛得出的前沿模型Elo排名来基准测试这些能力。通过发布此框架和锦标赛基础设施以供持续社区提交,我们建立了一个活性的、非饱和的测试平台,以持续衡量物理世界中开放式智能的扩展极限。更多信息请访问此URL。

英文摘要

To evaluate the frontier, we must measure models not by what they say, but by what they can engineer and build in grounded physical environments. We introduce \textsc{ArtifactArena}, an open-ended platform where models face a physically grounded hardware-software co-design challenge: engineering fully functional robots to compete in a simulated arena. We evaluate a frontier model's zero-shot, verifier guided refinement, and open-ended physical design capabilities through three harnesses that refine their bot artifacts based on text descriptions, physics simulator feedback, and gameplay data. We benchmark these capabilities with an Elo ranking of frontier models derived from head-to-head tournaments between their artifacts. By releasing this framework and tournament infrastructure for ongoing community submissions, we establish a living, non-saturating testbed to continuously measure the expanding limits of open-ended intelligence in the physical world. Please visit \href{https://artifactarena.ai}{https://artifactarena.ai} for more information.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑