发表机构
Nanyang Technological University; Shanghai Artificial Intelligence Laboratory(南洋理工大学; 上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PhysLDM提出统一潜扩散范式,结合时空VAE与扩散模型,实现高保真可变形体一次性仿真,在混沌动力学中优于回归方法,并支持零样本泛化与逆问题求解。
AI 中文摘要
高保真可变形体的神经仿真一直是计算机图形学和物理人工智能领域的基础性挑战。对于高分辨率三维体积网格的长时程预测是困难的:自回归方法容易受到误差累积的影响,而直接以原生分辨率进行多帧预测在计算上不可行。这促使我们探索一种紧凑的时空潜表示,而这一方向在基于网格的体积物理仿真中尚未得到充分研究。同时,确定性回归与生成式扩散哪一种预测范式更为合适仍不明确。为应对这些相互关联的挑战,我们提出了PhysLDM,一种用于一次性体积可变形仿真的统一潜扩散范式。其核心是一个整体性的时空变分自编码器,避免了标准时间压缩(如常见视频VAE)中的“阶梯”伪影,在米级场景中实现了约2.48毫米的重建精度,同时实现了高达78倍的token压缩。基于这一可靠的潜空间,我们系统地比较了回归和扩散方法。我们的实验揭示了一个关键的建模洞见:复杂的可变形动力学往往是混沌的,在这种情形下,确定性回归倾向于产生非物理的平均结果,而扩散方法能更好地建模其分布。因此,我们采用潜扩散模型,有效地从混沌数据中学习以生成物理上合理的轨迹。仅在一个Objaverse规模的数据集上进行纯运动学训练,单个PhysLDM即可零样本泛化到未见过的分布外数据集(GSO和Toys4K)。其可微性进一步支持高效求解逆问题和进行高阶设计优化。据我们所知,PhysLDM是首个用于体积可变形动力学的高保真时空自编码器和潜扩散范式,为神经仿真提供了一种可扩展且稳健的方法。
英文摘要
Neural simulation of high-fidelity deformable bodies is a foundational challenge in computer graphics and physical AI. Long-horizon prediction for high-resolution 3D volumetric meshes is difficult: autoregressive methods are susceptible to error accumulation, while direct multi-frame prediction at native resolution is computationally prohibitive. This motivates a compact spatiotemporal latent representation, which is largely unexplored for mesh-based volumetric physics. Meanwhile, it remains unclear whether deterministic regression or generative diffusion is the more appropriate predictive paradigm. To address these coupled challenges, we introduce PhysLDM, a unified latent-diffusion paradigm for one-shot volumetric deformable simulation. Its core is a holistic spatiotemporal VAE that avoids the "staircase" artifacts of standard temporal compression (as in common video VAEs), achieving ~2.48 mm reconstruction precision on meter-scale scenes at up to 78x token compression. Based on this reliable latent space, we systematically compare regression and diffusion methods. Our experiments uncover a key modeling insight: complex deformable dynamics are often chaotic, and in this regime deterministic regression tends to produce non-physical averages, whereas diffusion better models their distribution. Accordingly, we employ a latent diffusion model that effectively learns from the chaotic data to generate physically plausible trajectories. Trained purely kinematically on an Objaverse-scale dataset, a single PhysLDM generalizes zero-shot to unseen OOD datasets (GSO and Toys4K). Its differentiability further enables efficient solution of inverse problems and higher-order design optimization. To our knowledge, PhysLDM is the first high-fidelity spatiotemporal autoencoder and latent-diffusion paradigm for volumetric deformable dynamics, offering a scalable and robust approach to neural simulation.