分析动力学:从单目视频学习基于物理的表征以实现快速内在动力学推理
Analytic Dynamics: Learning Physics-Grounded Representation for Fast Intrinsic Dynamics Inference from Monocular Videos
浏览论文内容
中文总结 AI 辅助
针对从单目视频推断物体动力学的挑战,提出Analytic Dynamics框架,引入基于物理的中间表征,开发相关数据流水线与基准,实现高效准确可泛化的动力学推理。
中文摘要 AI 辅助
从视觉观测中推断物体动力学是智能体推理和与物理世界交互的基础,但由于视觉证据与内在动力学之间存在根本差距,该任务仍具挑战性。现有方法要么依赖昂贵的逐场景优化,限制了效率和可扩展性;要么直接将视觉证据映射到内在动力学,而不引入中间物理抽象,导致模型易受外观和几何捷径影响。为弥合这一差距,我们提出Analytic Dynamics,这是一种前馈动力学推理框架,在视觉观测和内在动力学之间引入基于物理的中间动力学表征。具体而言,我们利用模拟中可用的特权物理状态,包括位置、位移和变形梯度场,来学习仅从视觉观测中难以发现的结构化动力学表征。通过将视觉表征与该空间对齐,我们为视觉模型赋予基于物理的归纳偏置,引导其捕捉与动力学相关的模式,用于材料模型分类和参数回归。为推动该研究,我们开发了一个动力学数据生成流水线和基准,包含配对的物理状态轨迹、渲染视频以及真实材料模型和参数。大量实验表明,Analytic Dynamics可从单目视频实现高效、准确且可泛化的动力学推理。
英文摘要
Inferring object dynamics from visual observations is essential for intelligent agents to reason about and interact with the physical world, yet remains challenging due to the fundamental gap between visual evidence and intrinsic dynamics. Existing methods either rely on costly per-scene optimization, limiting efficiency and scalability, or directly map visual evidence to intrinsic dynamics without intermediate physical abstractions, making them prone to appearance and geometry shortcuts. To bridge this gap, we propose Analytic Dynamics, a feed-forward dynamics inference framework that introduces an intermediate physics-grounded dynamics representation between visual observations and intrinsic dynamics. Specifically, we leverage privileged physical states, including position, displacement, and deformation gradient fields, which are available in simulation, to learn a structured dynamics representation that is difficult to discover from visual observations alone. By aligning visual representations with this space, we equip visual models with a physics-grounded inductive bias, guiding them to capture dynamics-relevant patterns for material model classification and parameter regression. To facilitate this research, we develop a dynamics data generation pipeline and benchmark containing paired physical state trajectories, rendered videos, and ground-truth material models and parameters. Extensive experiments demonstrate that Analytic Dynamics achieves efficient, accurate, and generalizable dynamics inference from monocular videos.
发表机构
- School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。