ORMA:基于优化的单目四足动物4D重建
ORMA: Optimization-based Monocular 4D Reconstruction of Articulated Animals
- University of Cambridge(剑桥大学)
- Technical University of Munich(慕尼黑工业大学)
- Munich Center for Machine Learning(慕尼黑机器学习中心)
- IMATI-CNR(意大利国家研究委员会数学与应用技术研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
ORMA提出一种无需训练的单目4D重建框架,解耦关节与形状,利用生成式3D先验和SMAL+模型,结合DINO对应与时序一致性,在PAW4D等基准上提升重建精度与运动一致性。
AI中文摘要:
从单目视频中恢复动物的关节式4D表示仍然具有挑战性,这主要是因为四足动物形态的多样性以及缺乏动物4D监督数据。现有的基于学习的重建方法在单个图像上运行,并依赖于合成或模型拟合的3D监督,这继承了强参数先验的约束,并限制了对分布外物种的泛化能力。当应用于分布外的动物时,它们通常能恢复出合理的姿态,但产生不准确的几何形状,因为底层形状模型无法忠实地表示观察到的实例。我们提出了ORMA,一个无需训练的重建框架,它将关节运动与形状解耦,使用预测的姿态作为优化的参考,同时利用生成式3D先验进行准确的形状重建。给定参考图像,我们重建动物几何形状并将其注册到参数化模型SMAL+上,得到适应于观察实例的关节式形状。然后,我们将逐帧的关节姿态估计与全局一致的相机姿态相结合,在共享的世界坐标系中恢复动物运动,并使用自监督的DINO对应关系和时序一致性进一步细化重建。为了进行定量评估,我们引入了PAW4D,一个具有真实3D几何和相机运动的合成多物种基准。在PAW4D、PFERD和具有挑战性的野外视频上的实验表明,ORMA提高了重建准确性,同时恢复了跨不同四足物种的全局一致的动物运动。
英文摘要:
Recovering articulated 4D representations of animals from monocular videos remains challenging due to the large diversity of quadruped morphologies and lack of animal 4D supervision data. Existing learning-based reconstruction methods operate on individual images and rely on synthetic or model-fitted 3D supervision, which inherits the constraints of strong parametric priors and limits generalization to out-of-distribution species. When applied to out-of-distribution animals, they often recover a plausible pose while producing inaccurate geometry because the underlying shape model cannot faithfully represent the observed instance. We present ORMA, a training-free reconstruction framework that decouples articulation from shape, using the predicted pose as reference for optimization while leveraging generative 3D priors for accurate shape reconstruction. Given a reference image, we reconstruct the animal geometry and register it to the parametric model SMAL+, yielding an articulated shape adapted to the observed instance. We then combine per-frame articulated pose estimates with globally consistent camera poses to recover animal motion in a shared world coordinate frame, and further refine the reconstruction using self-supervised DINO correspondences and temporal consistency. To enable quantitative evaluation, we introduce PAW4D, a synthetic multi-species benchmark with ground-truth 3D geometry and camera motion. Experiments on PAW4D, PFERD, and challenging in-the-wild videos demonstrate that ORMA improves reconstruction accuracy while recovering globally consistend animal motion across diverse quadruped species.