arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09145cs.CV

Point4D:长程4D运动重建

Point4D: Long-range 4D Motion Reconstruction

Minsik Jeon, Jay Karhade, Deva Ramanan, Shubham Tulsiani

首次发表
浏览论文内容

中文总结 AI 辅助

Point4D是一种前馈模型,通过3D查询运动解码器和视觉描述符重用,实现数百帧长视频的密集4D轨迹重建,性能超越现有方法。

中文摘要 AI 辅助

我们提出了Point4D,一种用于长视频序列4D重建的前馈模型。与现有仅限于最多几十帧短输入窗口的4D方法不同,Point4D能够在跨越数百帧的视频中可靠地推断出密集的逐点3D轨迹。实现这一点的关键创新是我们灵活的基于3D查询的运动解码器,它将轨迹预测与图像平面可见性解耦。预测的3D端点随后直接在下一个块中重新查询,无需重新投影或匹配。此外,我们展示了从点可见的任意帧中提取并重用视觉描述符比仅依赖源补丁能带来更好的性能。总体而言,Point4D在跨越200多帧的多种长视频跟踪基准上达到了最先进的性能,并大幅超越了之前的前馈4D方法。项目页面:此https URL

英文摘要

We introduce Point4D, a feed-forward model for 4D reconstruction of long-range video sequences. Point4D is able to reliably infer dense per-point 3D trajectories across multi-hundred-frame videos, unlike existing 4D methods that are limited to short input windows of at most a few dozen frames. A key innovation that enables this is our flexible 3D query-based motion decoder that decouples trajectory prediction from image-plane visibility. The predicted 3D endpoints are then directly re-queried in the next chunk without re-projection or matching. Furthermore, we show that extracting and reusing a visual descriptor from an arbitrary frame where the point is visible leads to better performance than relying solely on the source patch. Overall, Point4D achieves state-of-the-art performance across diverse long-video tracking benchmarks spanning over 200 frames and largely outperforms previous feed-forward 4D method. Project page: https://point-4d.github.io

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

↑