arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GS-Agent:通过生成式模拟创建4D物理世界

GS-Agent: Creating 4D Physical Worlds With Generative Simulation

Hongxin Zhang, Chunru Lin, Junyan Li, Zhou Xian, Tsun-Hsuan Wang, Chuang Gan

arXiv 2607.21522首次发表:更新:

发表机构

University of Massachusetts Amherst; Genesis AI(马萨诸塞大学阿默斯特分校; 创世纪人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在从自然语言描述创建4D物理世界。核心方法是提出GS-Agent这一端到端多智能体框架,分解任务并让多智能体与物理引擎协作。主要贡献是能有效转换自然语言为物理合理的4D世界,实现相机和灯光控制,为新范式奠基。

AI 中文摘要

从自然语言描述创建动态且物理逼真的4D世界既迷人又具有挑战性。传统计算机图形方法依赖手动创建,需大量人力微调材料、运动和视觉保真度。生成基础模型的进展引发了从大规模数据生成此类4D世界的兴趣,但现有方法仍难以确保物理合理性和可控性。本文利用基础模型构建代理系统,提出端到端多智能体框架GS-Agent,将任务分解为实体管理和渲染配置,多智能体协作通过代码与物理引擎交互,迭代构建符合描述的4D世界。实验结果表明,GS-Agent能有效将自然语言转换为多样且物理合理的4D世界,实现电影级相机和灯光控制,为4D世界生成新范式奠定基础。

英文摘要

Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challenging. Traditional computer graphics methods rely on manual creation, requiring extensive human effort to fine-tune materials, motions, and visual fidelity. Recent advances in generative foundation models have sparked interest in learning to generate such 4D worlds from large-scale data; however, existing methods still struggle to ensure physical plausibility and controllability. In this work, we take a different path by leveraging foundation models to construct an agentic system that emulates how humans traditionally create 4D worlds, yet automates the entire process. We present GS-Agent, an end-to-end multi-agent framework that integrates physics engines in the loop to generate realistic, dynamic, and controllable 4D physical worlds from natural language. Inspired by how humans build 4D worlds, GS-Agent decomposes the task into entity management, covering 3D asset curation, material tuning, placement, and motion control, and rendering configuration, including camera and lighting manipulation. Multiple agents with distinct expertise interact with the physics engine via code, seek multimodal feedback, and collaborate to iteratively construct 4D worlds that align with the given descriptions. Experimental results show that GS-Agent effectively converts natural language into diverse and physically plausible 4D worlds exhibiting rich interactions among liquids, deformable objects, and rigid bodies, while achieving cinematic camera and lighting control. We envision GS-Agent as a foundation for a new paradigm in 4D world generation, empowering creative content creation and physical AI. Project page at https://umass-embodied-agi.github.io/gs-agent/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑