发表机构
The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出GCA框架,结合LoRA适配与RDGA、CPR模块,从单目动态视频学习隐式本构定律,在多类数据集上性能优于现有方法
AI 中文摘要
我们提出了GCA(Gaussian Constitutive Alignment,高斯本构对齐),这是一个从由3D高斯表示的可变形物体的单目动态视频中学习隐式本构定律的框架。给定用于几何初始化的静态多视图扫描,我们的方法仅从运动物体的单个固定视点视频中学习内在物理动力学。现有的隐式方法在噪声监督下常陷入局部极小值且缺乏物理解释性,而显式方法依赖预定义的本构方程,这限制了泛化能力并在单目设置中变得不稳定。为解决这些挑战,我们的框架将基于LoRA的适配与两个关键对齐模块相结合。首先,我们提出基于秩的深度-几何锚点(RDGA),通过尺度不变的基于秩的深度对齐从单目动态观测中建立鲁棒几何约束,减少对不可靠像素级颜色监督的依赖。其次,本构先验正则化器(CPR)将经典本构模型作为软可微分先验进行整合,在优化过程中进行正则化,同时保留隐式建模的灵活性——即使实际材料不在假设范围内也适用。在合成、真实到模拟及真实世界数据集上的大量实验表明,GCA的性能优于现有方法,在合成基准上比最强基线实现了低48%的 chamfer距离,同时在单目监督下保持鲁棒性。
英文摘要
We present GCA (Gaussian Constitutive Alignment), a framework for learning implicit constitutive laws from monocular dynamic video of deformable objects represented by 3D Gaussians. Given a static multi-view scan for geometric initialization, our method learns intrinsic physical dynamics solely from a single fixed-viewpoint video of the moving object. Existing implicit methods often suffer from local minima under noisy supervision and lack physical interpretability, while explicit approaches rely on predefined constitutive equations, limiting generalizability and becoming unstable in monocular settings. To address these challenges, our framework unifies LoRA-based adaptation with two key alignment modules. First, we propose Rank-based Depth-Geometric Anchors (RDGA) to establish robust geometric constraints from monocular dynamic observations via scale-invariant rank-based depth alignment, reducing the reliance on unreliable pixel-level color supervision. Second, a Constitutive Prior Regularizer (CPR) integrates classical constitutive models as soft differentiable priors, regularizing the optimization while preserving the flexibility of implicit modeling---even when the actual material is absent from the hypotheses. Extensive experiments on synthetic, real-to-sim, and real-world datasets demonstrate that GCA outperforms existing methods, achieving 48% lower Chamfer Distance than the strongest baseline on synthetic benchmarks while remaining robust under monocular supervision.