arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OmniX:通过前馈轨迹场进行任意视角和任意时刻的4D重建

OmniX: Any-view and Any-time 4D Reconstruction via Feed-forward Trajectory Fields

Yanqin Jiang, Tengfei Wang, Zhengwei Wang, Chenjie Cao, Junta Wu, Wenhan Luo, Weiming Hu, Jin Gao, Chunchao Guo

arXiv 2607.10840首次发表:更新:

发表机构

Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences; Tencent Hunyuan; HKUST; Beijing Key Laboratory of Super Intelligent Security of Multi-Modal Information; School of Information Science and Technology, ShanghaiTech University(中国科学院自动化研究所; 中国科学院大学; 腾讯混元; 香港科技大学; 多模态信息超智能安全北京重点实验室; 上海科技大学信息科学与技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对前馈4D重建方法局限,提出OmniX框架,通过解耦动态运动建模与静态几何预测,利用动态令牌和3D运动结构生成轨迹场,并构建数据引擎和数据集,在多任务上取得领先性能。

AI 中文摘要

以往的前馈4D重建方法要么预测每帧静态点云而忽略前景运动,要么估计点云轨迹但限于小相机运动,无法在大视角变化下重建完整动态场景。为此提出OmniX,它能从大相机运动视频中为每个像素预测密集3D点轨迹。该方法将动态运动建模与静态几何预测解耦,用紧凑动态令牌表示运动,利用3D运动的稀疏和低秩结构生成轨迹场。还构建自动UE5 4D数据引擎并引入大规模数据集。OmniX在密集3D点轨迹预测等任务上取得了领先性能。

英文摘要

Previous feed-forward 4D reconstruction methods either predict per-frame static point clouds, ignoring foreground motion, or estimate point cloud trajectories while being limited to small camera motions. This restricts their ability to aggregate observations over time and reconstruct complete dynamic scenes under large viewpoint changes. To address this limitation, we propose OmniX, a feed-forward 4D reconstruction framework that predicts dense 3D point trajectories for every pixel from videos with large camera motion. OmniX decouples dynamic motion modeling from static geometry prediction and represents motion using a compact set of dynamic tokens. By leveraging the sparse and low-rank structure of 3D motion, these tokens generate trajectory fields for all pixels across all images while efficiently preserving global interactions. To facilitate training, we further build an automatic UE5-based 4D data engine and introduce a large-scale dataset containing 80K scenes and 1.28M multi-view videos with full geometric annotations. OmniX achieves state-of-the-art performance on dense 3D point trajectory prediction and 3D point tracking, while also demonstrating competitive results on video depth estimation and camera pose estimation.

CommentsAccepted by ECCV 2026, project page: https://omnix4d.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑