arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MotionForesight:重新利用视频模型进行未来3D场景流预测

MotionForesight: Re-purposing Video Models for Future 3D Scene-Flow Prediction

Homanga Bharadhwaj, Yash Jangir

arXiv 2607.16192首次发表:更新:

AI 中文总结

研究如何从人类与物体交互的单目视频中学习预测物体未来3D轨迹,核心方法是利用预训练视频模型,通过特定操作训练预测器,仅用少量视频且无辅助输入就能泛化,优于大模型,为具身智能提供有效几何预测。

AI 中文摘要

人类能从被动观察中推断物体的可能运动,本文研究如何从人类与物体交互的普通单目视频中学习这种预测。MotionForesight在给定短观察视频上下文时,预测被操作物体上点的未来3D轨迹,将交互预测视为以物体为中心的3D运动预测。通过从基于预训练视频模型构建的密集3D跟踪器开始,生成伪地面真值轨迹并仅用观察帧训练预测器,还使用学习到的掩码潜变量替换未来RGB和几何信息并训练轻量级适配器。仅用40k人类视频且无语言等辅助输入,该方法能在不同分布的物体、环境、观点和交互中泛化,且优于使用超百万训练视频的更大模型。这些结果表明可将视频先验有效地重新用于具身智能的显式几何预测。

英文摘要

Humans can infer how objects are likely to move from passive observation: a cup may be lifted, a drawer may slide, and a lid may rotate shut. Such predictions expose the physical consequences of interaction needed to act in the real world. We study how to learn this anticipation from ordinary monocular videos of human-object interaction. Given a short observed video context, MotionForesight predicts future 3D trajectories for points on the manipulated object. This casts interaction prediction as object-centered 3D motion forecasting without any assumptions on the object properties. Our key insight is that video prediction models already encode rich priors about how objects move during human interactions. We redirect these priors from pixel prediction toward future 3D scene flow. We start from a dense 3D tracker built on a pretrained video model, generate pseudo-ground-truth tracks from complete clips, and train the forecaster using only the observed frames. We replace future RGB and geometry with learned mask latents and train a lightweight adapter to turn the retrospective tracking representation into a forward predictor, while freezing the large video and tracking components. Using just 40k human videos and no auxiliary inputs such as language, MotionForesight generalizes across diverse out-of-distribution objects, environments, viewpoints, and interactions. It also outperforms substantially larger models that use over a million training videos. These results show that we can efficiently re-purpose video priors into explicit geometric forecasts for embodied intelligence. https://motionforesight.github.io/

CommentsWebsite with visual results motionforesight.github.io

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑