arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SM4RT:学习用于4D重建的结构化运动几何

SM4RT: Learning Structured Motion Geometry for 4D Reconstruction

Shing Ho J. Lin, Wenzhao Zheng, Dong Zhuo, Yuqi Wu, Jie Zhou, Jiwen Lu

arXiv 2607.22534首次发表:更新:

发表机构

Intelligent Vision Group, Tsinghua University(清华大学智能视觉组)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对将单目3D重建能力扩展到4D动态理解的挑战,提出SM4RT,通过引入运动结构表示场景动态,利用并行运动几何编码器和解码器,从单目RGB视频中联合推断相关信息,实现强大运动重建性能并保留场景运动几何结构。

AI 中文摘要

几何基础模型(GFMs)极大地推动了单目3D重建,但将此能力扩展到4D动态理解仍是一项重大挑战。现有多数运动感知方法将运动视为独立的逐点位移,忽略了物理运动的结构化本质。而实际物体通常遵循刚体运动学规律,点通常集体移动而非孤立移动。基于此,我们提出了SM4RT,一种用于端到端3D重建和结构化运动感知的结构化运动4D重建变换器。SM4RT引入运动结构来表示场景动态,将场景运动分解为一组紧凑的运动基,每个运动基表示为SE(3)中6D扭转的时间序列。然后通过对这些基的稀疏、逐像素时间共享分配权重来恢复密集场景运动。SM4RT引入了并行运动几何编码器和解码器,可从单目RGB视频中单次前向传递联合推断3D几何、世界坐标运动和场景运动学结构。SM4RT在保留场景运动几何结构的同时实现了强大的运动重建性能。

英文摘要

Geometry Foundation Models (GFMs) have substantially advanced monocular 3D reconstruction, yet extending this capability to 4D dynamic understanding remains a fundamental challenge. Most existing motion perception methods (e.g., sparse tracking, dense point-wise flow) treat motion as independent point-wise displacements, ignoring the structured nature of physical motion. However, real-world objects usually obey rigid-body kinematics, and points thus usually move collectively, not in isolation. Motion itself possesses geometric structure: physical objects undergo a set of rigid-body transformations governed by SE(3), rather than unstructured point-wise displacements. Building on this insight, we propose SM4RT, a Structured Motion 4D Reconstruction Transformer for end-to-end 3D reconstruction and structured motion perception. SM4RT introduces Structure-of-Motion to represent scene dynamics, where scene motion is decomposed into a compact set of motion bases, each represented as a temporal sequence of 6D twists in SE(3). Dense scene motion is then recovered by sparse, time-shared per-pixel assignment weights over these bases, ensuring points on the same object share a common rigid-body motion trajectory. SM4RT introduces a parallel motion geometry encoder and decoder that jointly infer 3D geometry, world-coordinate motion, and scene kinematic structure in a single forward pass from monocular RGB video. SM4RT achieves strong motion reconstruction performance while preserving the geometric structure of scene motion.

CommentsCode is available at: https://github.com/wzzheng/SM4RT

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑