PhysMAS:基于物理的多智能体组合式4D高斯合成
PhysMAS: Physics-Grounded Multi-Agent Synthesis of Compositional 4D Gaussians
浏览论文内容
中文总结 AI 辅助
PhysMAS提出基于物理的多智能体框架,通过物体-部件场景智能体和材料推理智能体实现异构多部件及交互多物体场景的4D高斯合成,无需SDS反向传播,提升语义对齐和物理合理性并减少运行时间。
中文摘要 AI 辅助
高效、全自动且物理上合理的4D高斯合成是动态场景生成的重要目标。最近的基于物理的方法将3D高斯与物质点法(MPM)耦合以生成物理驱动的运动,但将该范式扩展到异构多部件物体和相互交互的多物体场景仍然具有挑战性。物体级物理分配将不同部件折叠为单一材料状态,而来自大型语言模型、视觉语言模型或智能体的一次性预测既不能可靠地将不同材料绑定到已识别的部件,也不能验证所得的MPM配置是否可执行。同时,基于分数蒸馏采样(SDS)的参数优化需要对每个场景进行重复的分数评估和梯度反向传播,导致优化时间过长,并可能产生次优或不稳定的解。因此,我们提出了PhysMAS,一个基于物理的多智能体框架。根据运动提示和四个场景视图,物体-部件场景智能体建立持久身份,并调用材料推理智能体获取部件级配置。它调用求解器感知技能将这些身份和配置绑定到每个粒子的MPM场,并在共享域中执行所有物体;然后框架筛选候选的前向模拟结果。这支持异构多部件和交互式多物体场景,而无需逐场景扩散分数反向传播。大量实验表明,与依赖SDS的近期基于物理的4D高斯基线相比,PhysMAS实现了更好的语义对齐和感知物理合理性,同时所需运行时间更少。
英文摘要
Efficient, fully automatic, and physically plausible 4D Gaussian synthesis is an important goal for dynamic scene generation. Recent physics-based methods couple 3D Gaussians with the Material Point Method (MPM) to generate physically driven motion, but extending this paradigm to heterogeneous multi-part objects and interacting multi-object scenes remains challenging. Object-level physical assignment collapses distinct parts into a single material state, while one-shot predictions from large language models, vision-language models, or agents neither reliably bind different materials to identified parts nor verify that the resulting MPM configuration is executable. Score Distillation Sampling (SDS)-based parameter optimization, meanwhile, requires repeated per-scene score evaluations and gradient backpropagation, incurring lengthy optimization and potentially yielding suboptimal or unstable solutions. We therefore present PhysMAS, a physics-grounded multi-agent framework. From a motion prompt and four scene views, an Object-Part Scene Agent establishes persistent identities and calls a Material Reasoning Agent for part-wise profiles. It invokes solver-aware skills to bind these identities and profiles to per-particle MPM fields and execute all objects in a shared domain; the framework then screens candidate forward-simulation results. This supports heterogeneous multi-part and interacting multi-object scenes without per-scene diffusion-score backpropagation. Extensive experiments demonstrate that, compared with recent physics-based 4D Gaussian baselines that rely on SDS, PhysMAS achieves better semantic alignment and perceived physical plausibility while requiring less runtime.
发表机构
- Beijing Institute of Technology(北京理工大学)
- Meituan(美团)
- Sichuan University(四川大学)
机构由 AI 辅助整理,请以论文原文为准。