arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LMBuild:评估LLM智能体生成可构建且功能完备的结构

LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures

Jiateng Liu, Rushi Wang, Cheng Qian, Xuejun Zhang, Sun Li, Jiayu Liu, Yifan Shen, Xu Cao, Jiarui Yao, Bingxuan Li, Ruhi Sarikaya, Heng Ji

arXiv 2610.04292首次发表:更新:

发表机构

University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出LMBuild基准,评估LLM智能体生成可构建且功能完备结构的能力,发现功能可供性与物理可操作性仍是主要挑战,提供功能规格可显著提升性能。

AI 中文摘要

基于LLM的智能体生成复杂三维结构的能力日益增强,有望重塑物体在物理世界中的设计与实现方式。然而,生成优美的几何形状与生成可构建并能执行预期功能的物体有着根本区别。现有评估大多聚焦于几何质量,而忽视了物理可实现性。我们提出LMBuild,一个用于评估LLM智能体生成可构建且功能完备结构的基准。LMBuild将生成的对象表示为包含部件分解、关节、材料和序列的装配结构。为支持可复现的评估,我们提供了一个统一框架,包括:(1)一个交互式环境,智能体可在其中使用工具检索、创建和放置组件以构建对象;(2)一个精选基准,重新利用已有的CAD数据集,并利用维基百科的知识对其进行增强;(3)一个评估框架,涵盖结构稳健性、功能可供性、设计质量和物理实现。对30个系统的评估揭示了几个有趣的发现:(a)稳健性和对齐不再是前沿闭源模型的主要瓶颈,而功能可供性和物理可操作性仍然更具挑战性;(b)更强的模型能更有效地创建新组件,而较弱的模型倾向于依赖检索;(c)提供功能规格说明能显著提高部件完整性、运动学和物理可操作性。这些结果表明,生成真实世界结构需要对功能可供性、力学以及设计和创建新颖组件进行更深入的推理。我们期望LMBuild能为衡量进展和激励研究提供基础,以推动生成可构建且功能完备结构的智能体。

英文摘要

LLM-based agents are increasingly capable of generating complex 3D structures, with the potential to reshape how objects are designed and realized in the physical world. Yet, producing elegant geometry is fundamentally different from producing objects that can be built and perform their intended functions. Existing evaluations largely focus on geometric quality while overlooking physical realizability. We introduce LMBuild, a benchmark for evaluating LLM agents on generating buildable and functional structures. LMBuild represents generated objects as assembled structures comprising part decompositions, joints, materials, and sequences. To support reproducible evaluation, we provide a unified framework consisting of: (1) an interactive environment in which agents can use tools to retrieve, create, and place components to construct objects; (2) a curated benchmark that repurposes established CAD datasets and augments them with knowledge from Wikipedia; and (3) a evaluation framework covering structural soundness, functional affordance, design quality, and physical realization. Evaluations across 30 systems reveal several intriguing findings: (a) Soundness and alignment are no longer the primary bottlenecks for frontier closed-source models, while functional affordance and physical operability remain substantially more challenging; (b) stronger models more effectively create new components, whereas weaker models tend to rely on retrieval; and (c) providing functional specifications substantially improves part completeness, kinematics, and physical operability. These results show that generating real-world structures requires deeper reasoning about functional affordances, mechanics, and designing and creating novel components. We expect LMBuild to provide a foundation for measuring progress and incentivizing research toward agents that generate buildable and functional structures.

Comments64 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑