arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LEGO-Anything:用于3D场景重建的编码智能体

LEGO-Anything: Coding Agents for 3D Scene Reconstruction

Xirui Li, Peng Shi, Mingwen Dong, Sheng Zhang, Zhuoyan Xu, Dongkyu Lee, Shuaichen Chang, Yi Xiang, Lin Pan, Jiarong Jiang

arXiv 2609.36380首次发表:更新:

发表机构

University of Maryland, College Park; AWS(马里兰大学学院公园分校; 亚马逊云服务)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出LEGO-Anything框架,利用编码智能体迭代编写Blender代码重建3D场景,并引入LEGO-Bench基准评估,发现现有智能体在场景重建中的不足,进而提出LEGO-Plugin插件以提升性能。

AI 中文摘要

从单张图像重建的3D场景,当其被表示为显式场景程序(而非渲染结果或固定的3D输出)时最为有用,因为执行该程序可生成可检查、可编辑和可查询的场景。我们提出LEGO-Anything,一个图像到代码(Image-to-Code)框架,其中编码智能体迭代地编写并执行Blender代码,检查场景和渲染结果,并修订程序。为了评估端到端的场景恢复,我们引入LEGO-Bench,一个基于模拟器的基准,包含来自104个多样化室内和室外场景的208张图像。LEGO-Bench分别对工件有效性、可见表面几何和渲染外观进行评分。其基于模拟器的设计支持可扩展性和精确的自动评估。在评估的智能体中,GPT-6-astra取得了最强的总体结果,室内得分为53.4%,室外得分为39.6%,但在生成有效场景工件与忠实恢复场景几何和外观之间仍存在显著差距。对智能体构建轨迹的分析揭示了三个反复出现的问题:场景初始化薄弱、迭代过程中的回归性编辑以及不可靠的自我评估。这些发现促成了LEGO-Plugin,一个无需训练的插件框架,用于更可控的迭代场景构建,它改进了所有六个评估模型,总体得分相对提升高达62.7%。最后,我们测试重建场景是否能表示自然图像并支持视觉任务。在LEGO-World中,我们将对象检测、实例掩码和相对深度作为对GPT-6-astra重建场景的确定性查询。这些读出结果在所有三个任务上显示出非平凡的性能,但远不及专门的视觉模型,这表明当前编码智能体构建的程序化场景是有前景的,但尚不足以精确表示自然图像。

英文摘要

A 3D scene reconstructed from a single image is most useful when represented not as a rendering or a fixed 3D output, but as an explicit scene program whose execution yields a scene that can be inspected, edited, and queried. We present LEGO-Anything, an Image-to-Code framework in which a coding agent iteratively writes and executes Blender code, inspects scenes and renderings, and revises the program. To evaluate end-to-end scene recovery, we introduce LEGO-Bench, a simulator-grounded benchmark with 208 images from 104 diverse indoor and outdoor scenes. LEGO-Bench separately scores artifact validity, visible-surface geometry, and rendered appearance. Its simulator-grounded design enables extensibility and precise automatic evaluation. Among evaluated agents, GPT-6-astra achieves the strongest overall results, with 53.4% indoor and 39.6% outdoor scores, yet substantial gaps remain between delivering valid scene artifacts and faithfully recovering scene geometry and appearance. Analysis of agent construction trajectories reveals three recurring issues: weak scene initialization, regressive edits during iteration, and unreliable self-evaluation. These findings motivate LEGO-Plugin, a training-free harness plugin for more controlled iterative scene construction, which improves all six evaluated models, with relative gains of up to 62.7% in overall score. Finally, we test whether reconstructed scenes can represent natural images and support vision tasks. In LEGO-World, we derive object detections, instance masks, and relative depth as deterministic queries on scenes reconstructed by GPT-6-astra. These readouts show non-trivial performance across all three tasks but fall well short of specialized vision models, suggesting that program-constructed scenes from current coding agents are a promising but not yet sufficiently precise representation of natural images.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑