arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36777cs.AI

Code4Scene:用于构建和编辑3D场景的编码智能体基准

Code4Scene: Benchmarking Coding Agents for Constructing and Editing 3D Scenes

Xiaokang Ye, Siddhant Hitesh Mantri, Zimeng Chen, Edward Zhang, Zhaoxu Zheng, Yuanheng Li, Yizhao Chen, Tianyang Huang, Lianhui Qin

首次发表
浏览论文内容

中文总结 AI 辅助

提出Code4Scene基准,含190个Unreal Engine案例,评估编码智能体在构建和编辑3D场景中的空间推理与状态控制,发现生成与可靠控制间存在显著差距。

中文摘要 AI 辅助

前沿编码智能体现在能够编写并执行代码来生成3D环境,但它们是否可靠地理解3D结构并精确控制场景状态仍不清楚。生成的3D场景是一个持久的、可执行的工件:一个令人信服的渲染可能隐藏错误的空间关系、相交物体或非预期的修改。我们引入了Code4Scene,一个由190个Unreal Engine案例组成的基准,这些案例基于人工组装的场景构建,在共享执行接口下评估编码智能体的两个互补设置。构建测试从开放式语言规范中进行场景级空间推理,其中许多实现都是有效的;编辑测试对场景状态进行精确控制,智能体必须从参考图像中恢复目标场景,同时保留其他一切内容。Code4Scene不是对代码或渲染视图进行评分,而是评估生成的引擎原生场景的任务完成度、工件完整性和静态物理有效性,编辑结果还会与保留的真实值进行比较。在95个案例的公开集上,跨14种编码智能体配置,构建和编辑性能强相关但不可互换(Spearman ρ = 0.78):Claude Fable 5.1在构建方面领先,Gemini 3.8 Flash在编辑方面领先,而GPT-6 Astra在总体上略微领先。空间组合是每个智能体最弱的构建类别,而编辑仍然不精确:最佳的修复F1仅为0.527,且35.8%的完全恢复目标的编辑仍在场景其他部分引入了非预期更改。这些结果揭示了看似合理的3D生成与可靠的空间推理和状态控制之间的差距。

英文摘要

Frontier coding agents can now write and execute code that authors 3D environments, but whether they reliably understand 3D structure and precisely control scene state remains unclear. The generated 3D scene is a persistent, executable artifact: a convincing render can hide incorrect spatial relations, intersecting objects, or unintended modifications. We introduce Code4Scene, a benchmark of 190 Unreal Engine cases built from human-assembled scenes that evaluates coding agents on two complementary settings under a shared execution interface. Construction tests scene-level spatial reasoning from open-ended language specifications, where many realizations are valid; editing tests precise control of scene state, where the agent must recover the target scene from reference images while preserving everything else. Rather than scoring code or rendered views, Code4Scene evaluates the generated engine-native scene for task fulfillment, artifact integrity, and static physical validity, with edits additionally compared against withheld ground truth. Across 14 coding-agent configurations on the 95-case public set, construction and editing performance are strongly correlated but not interchangeable (Spearman $ρ= 0.78$): Claude Fable 5.1 leads construction, Gemini 3.8 Flash leads editing, and GPT-6 Astra narrowly leads overall. Spatial Composition is the weakest construction category for every agent, while editing remains imprecise: the best Repair F1 is only 0.527, and 35.8% of edits that fully recover the target still introduce unintended changes elsewhere in the scene. These results expose a gap between plausible 3D generation and reliable spatial reasoning and state control.

发表机构

  • UC San Diego(加州大学圣迭戈分校)
  • UC Berkeley(加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑