arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MoSE3:学习每个像素的世界空间SE(3)

MoSE3: Learning World-Space SE(3) at Every Pixel

Jiahuan Cheng, Zhiyi Li, Tian Xia, Ruojin Cai, Yilun Du, Qianqian Wang

arXiv 2610.03716首次发表:更新:

发表机构

Harvard University; Kempner Institute; Johns Hopkins University; MIT(哈佛大学; 肯普纳研究所; 约翰斯·霍普金斯大学; 麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MoSE3是首个从单目视频逐像素预测世界空间SE(3)运动的前馈模型,通过3D点轨迹和刚性嵌入联合学习,在刚性和铰接基准上达到最先进性能,并泛化至真实视频。

AI 中文摘要

稠密3D点跟踪一直是动态场景运动建模的重要范式,但点轨迹仅仅是每个像素的3自由度平移曲线:它捕捉像素的去向,却不包含底层部件的旋转,也不包含哪些像素作为一个整体一起运动。我们提出MoSE3,这是首个从前馈模型,能够从单目RGB视频预测稠密SE(3)运动,在世界空间中为每个像素生成完整的6自由度刚体变换。逐像素SE(3)运动提供了对场景运动更丰富的视角:旋转、平移和分组同时呈现。直接预测SE(3)具有挑战性:旋转位于弯曲流形上,不适合欧几里得回归,且SE(3)标注尤其难以获取。为应对这些挑战,MoSE3通过两个联合学习的中间表示——3D点轨迹和刚性嵌入——来预测逐像素SE(3),并通过在每个软刚性簇内可微地拟合变换来恢复SE(3),从而实现端到端预测和监督。为弥补数据缺口,我们引入Art-Kubric,一个大规模合成数据集,包含具有丰富物理交互的铰接物体的稠密SE(3)和刚性标签。MoSE3在刚性和铰接基准上均达到了像素、部件和物体级别的SE(3)估计最先进性能,并在三个数据集上取得了平均3D点跟踪精度的最先进水平,同时尽管仅使用合成运动数据训练,仍展现出对真实世界视频的强泛化能力。

英文摘要

Dense 3D point tracking has been a prominent paradigm for modeling motion in dynamic scenes, but a point track is just a 3-DoF translation curve per pixel: it captures where pixels go, not the rotation of the underlying part, nor which pixels move together as one body. We propose MoSE3, the first feed-forward model that predicts dense SE(3) motion from monocular RGB video, producing full 6-DoF rigid transforms at every pixel in world space. Per-pixel SE(3) motion offers a richer view of how a scene moves: rotation, translation, and grouping all at once. Directly predicting SE(3) is challenging: rotations lie on a curved manifold that is ill-suited to Euclidean regression, and annotations for SE(3) are particularly difficult to acquire. To address these challenges, MoSE3 predicts per-pixel SE(3) through two jointly learned intermediates, 3D point tracks and rigidity embeddings, and recovers SE(3) by differentiably fitting transforms within each soft rigid cluster, enabling end-to-end prediction and supervision. To close the data gap, we introduce Art-Kubric, a large-scale synthetic dataset with dense SE(3) and rigidity labels for articulated objects with rich physical interactions. MoSE3 achieves state-of-the-art SE(3) estimation at pixel, part, and object levels on both rigid and articulated benchmarks, and state-of-the-art average 3D point tracking accuracy across three datasets, while showing strong generalization to real-world videos despite being trained solely on synthetic motion data.

CommentsNeurIPS 2026 Spotlight. Project page: https://mose3-tracker.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑