arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33208cs.AIcs.CV

WorldAgent: 验证引导的智能体物理世界构建

WorldAgent: Verification-Guided Agentic Physical World Construction

Caoliwen Wang, Mengdi Wang, Yige Chen, Zejia Wu, Bowen Huang, Siyuan Chen, Guanxiong Chen, Lifu Wei, Heng Zhang, Qinghai Zhang, Yin Yang, Guandao Yang, Shiying … 展开作者

Caoliwen Wang, Mengdi Wang, Yige Chen, Zejia Wu, Bowen Huang, Siyuan Chen, Guanxiong Chen, Lifu Wei, Heng Zhang, Qinghai Zhang, Yin Yang, Guandao Yang, Shiying Xiong, Peng Wang, Chenfanfu Jiang, Peter Yichen Chen

中文总结 AI 辅助

WorldAgent是一个从自然语言提示构建物理世界的智能体框架,通过验证引导自动修订,无需用户调试,在AgenticSimBench上五项指标最优,用户研究评分最高。

中文摘要 AI 辅助

从语言构建复杂的物理世界需要协调广泛的3D环境、不同空间尺度下的详细结构和物体,以及在既定目标和隐含物理约束下相互作用的物理过程。我们提出WorldAgent,一个智能体框架,用于从单个自然语言提示进行验证引导的物理世界构建,无需迭代的用户调试。世界构建层将提示扩展为结构化的世界规范,并利用物理知识构建场景和运行数值模拟。在每一步之后,验证层检查场景几何和模拟状态以及渲染视图。失败的检查引导对规范的自动修订和相关步骤的重新执行。通过验收的世界通过所需检查,并保持可编辑以供进一步检查和重新模拟。我们引入AgenticSimBench,在该基准上,WorldAgent在评估的基于智能体的方法中,在七项指标中的五项上取得了最佳分数。在一项有26名参与者的用户研究中,它在所有四个标准上获得了最高的平均评分。

英文摘要

Constructing complex physical worlds from language requires coordinating extensive 3D environments, detailed structures and objects at different spatial scales, and interacting physical processes under both stated goals and implicit physical constraints. We present WorldAgent, an agentic framework for verification-guided physical world construction from a single natural-language prompt, without iterative user debugging. A world construction layer expands the prompt into a structured world specification and uses physical knowledge to build scenes and run numerical simulations. After every step, a verification layer inspects scene geometry and simulation states alongside rendered views. Failed checks guide automatic revisions to the specification and re-execution of the affected steps. Accepted worlds pass the required checks and remain editable for further inspection and resimulation. We introduce AgenticSimBench, on which WorldAgent achieves the best scores among the evaluated agent-based methods on five of seven metrics. In a 26-participant user study, it receives the highest mean ratings across all four criteria.

↑