arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18498cs.CV

DyG²T:基于3D高斯时空粒子图Transformer的对象动力学建模

DyG$^2$T: Modeling Object Dynamics with 3D Gaussian Temporal-Spatial Particle Graph Transformer

  • School of Computer Science and Technology, Harbin Institute of Technology at Weihai(哈尔滨工业大学(威海)计算机科学与技术学院)
  • Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
  • School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

Yansong Wang, Zhaobo Qi, Xinyan Liu, Beichen Zhang, Shuhui Wang, Weigang Zhang, Qingming Huang

AI总结:

本文针对现有动力学建模方法丢失细粒度细节、轨迹漂移等问题,提出DyG²T框架,通过空间补全关键点、TDN增强时间判别性、粒子图Transformer建模交互,在合成与真实数据集上实现精准动力学建模,泛化能力强。

AI中文摘要:

从有限视觉观测中建模对象动力学是实现具身交互场景中精准运动轨迹预测的基础问题。现有动力学建模方法先将重构的粒子表示压缩为稀疏关键点,再通过局部约束交互建模其演化,会丢失细粒度局部细节,模糊跨时空尺度的判别式交互建模,导致轨迹漂移和外观预测不准确。为解决这些问题,本文提出DyG²T,这一动力学建模框架通过空间补全、时间判别关键点表示,并在粒子图上建模多尺度交互来推断对象运动轨迹。空间上,DyG²T通过聚合相邻原始粒子位置丰富每个关键点,以恢复细粒度局部细节,同时显式编码关键点间的相对偏移以增强几何结构感知。时间上,本文引入时间解耦网络(TDN)识别隐空间中主导的跨帧变化并放大帧间差异,生成时间判别式表示,随后通过时间注意力聚合以捕捉逐帧时间演化线索。为实现全面的交互建模,粒子图Transformer利用全局注意力保留关键点间的判别式长程依赖,缓解局部约束建模导致的表示同质化,为精准轨迹预测提供稳健基础。在合成和真实世界数据集上的实验表明,DyG²T实现了精准的动力学建模与推理,并展现出强跨对象及真实世界泛化能力。

英文摘要:

Modeling object dynamics from limited visual observations is a fundamental problem for enabling accurate motion trajectory prediction in embodied interaction scenarios. Existing dynamics modeling methods first compress reconstructed particle representations into sparse Key Points and model their evolution using locally constrained interactions, thereby discarding fine-grained local details and obscuring discriminative interaction modeling across spatial and temporal scales, leading to drifting trajectories and inaccurate appearance prediction. To tackle these issues, we propose DyG$^2$T, a dynamics modeling framework that infers object motion trajectories by spatially completing and temporally discriminating Key Point representations and modeling multi-scale interaction over particle graphs. Spatially, DyG$^2$T enriches each Key Point by aggregating neighboring raw particle positions to recover fine-grained local details, while explicitly encoding relative offsets among Key Points to enhance geometric structure perception. Temporally, we introduce a Temporal Disentangling Network (TDN) to identify dominant cross-frame variations in latent space and amplify inter-frame differences, yielding temporally discriminative representations that are subsequently aggregated via Temporal Attention to capture frame-wise temporal evolution cues. For comprehensive interaction modeling, a Particle Graph Transformer leverages global attention to preserve discriminative long-range dependencies among Key Points, mitigating representation homogenization induced by locality-constrained modeling and providing a robust basis for accurate trajectory prediction. Experiments on both synthetic and real-world datasets demonstrate that DyG$^2$T achieves accurate dynamics modeling and reasoning, and exhibits strong cross-object and real-world generalization.

↑