arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12452cs.AIcs.CVcs.GR

BrickBench:评估智能体的乐高套装设计能力

BrickBench: Evaluating Agentic Brick Design

Peter Kulits, Yiqing Xu, R. Kenny Jones, Cordelia Schmid, Jiajun Wu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出BrickBench基准与BrickAgent环境,评估智能体的文本条件乐高设计能力,发现主流智能体虽满足部分要求但不及人类设计,相关资源已公开。

中文摘要 AI 辅助

我们提出了BrickBench,这是一个用于评估智能体在文本条件下乐高套装设计能力的基准。给定一个提示,智能体的任务是生成一个组装方案,该方案不仅要满足语义和设计标准,还要能够实际搭建。为此,它必须从离散零件库中选择零件,并对局部和全局约束进行联合推理。我们在三个规模和零件可用性各不相同的设置中,对有效性、对齐性和设计质量进行评分。我们提供了BrickAgent,这是一个供编码智能体构建、检查和验证其设计的环境。我们发现,主流智能体在很大程度上满足了可验证的物理和语义要求,但仍不及人类设计的水平。我们在该http链接发布了我们的基准和环境。

英文摘要

We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built. To do so, it must select parts from a discrete library and reason jointly about local and global constraints. We score validity, alignment, and design across three settings that vary in scale and part availability. We provide BrickAgent, an environment for coding agents to construct, inspect, and validate their designs. We find that leading agents largely satisfy verifiable physical and semantic requirements, but fall short of human designs. We release our benchmark and environment at http://www.brickben.ch

发表机构

  • Stanford University(斯坦福大学)
  • Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
  • Inria(法国国家信息与自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑