arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

30000小时自我中心视频无法教会的东西

What 30,000 Hours of Ego-centric Video Does Not Teach

Jiahua Dong, Anurag Bagchi, Yash Jangir, Muhammad Zubair Irshad, Sergey Zakharov, Martial Hebert, Homanga Bharadhwaj, Yu-Xiong Wang, Vitor Campagnolo Guizilini, Pavel Tokmakov

arXiv 2610.12464首次发表:更新:

发表机构

University of Illinois Urbana-Champaign; Carnegie Mellon University; Johns Hopkins University; Toyota Research Institute(伊利诺伊大学厄巴纳-香槟分校; 卡内基梅隆大学; 约翰斯·霍普金斯大学; 丰田研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究利用30000小时自我中心视频数据集,发现扩展数据可提升智能体建模但对象保真度提升缓慢,通过视觉条件设计和监督方案优化,表明缩小建模差距需关注训练方式而非仅数据量。

AI 中文摘要

世界模型为基于物理的模拟器提供了一种很有前景的替代方案,但仍远未达到实际部署的程度。我们利用包含1000多种场景类型和14000名贡献者的数据集,探究自我中心人类视频的扩展能将世界模型推进到何种程度。我们不依赖不透明的下游指标,而是在具有挑战性的分布外基准上直接评估智能体和对象交互的保真度。将训练数据增加100倍可同时提升两者,但提升程度不均:智能体建模效果良好,而对象保真度仍低得多且提升缓慢。我们证明智能体性能提升不一定来自数据,精心设计的视觉条件仅用一小部分数据就能使保真度达到饱和,这让我们可以单独测量对象保真度并发现其饱和点。随后,我们引入一种监督方案,将模型容量从场景外观转向对象动态,从而提升了对象保真度,但仍存在显著差距。最后,我们的结论可迁移到下游人形机器人建模中。总体而言,我们的结果表明,扩展自我中心数据使智能体建模接近其极限,而其对世界建模的影响仍远远落后,缩小这一差距将取决于模型的训练方式,而非仅取决于数据量。

英文摘要

World models offer a promising alternative to physics-based simulators, yet remain far from practical deployment. We ask how far scaling ego-centric human video takes them, using a dataset of 30,000 hours spanning over 1,000 scene types and 14,000 contributors. Rather than relying on opaque downstream metrics, we directly evaluate agent and object-interaction fidelity on a challenging out-of-distribution benchmark. Increasing training data by 100x improves both, but unevenly: the agent is modeled well, while object fidelity remains far lower and improves slowly. We show that the agent gains need not come from data, and a careful visual conditioning design saturates fidelity with a fraction of it, which lets us measure object fidelity on its own and discover its saturation point. We then introduce a supervision scheme that shifts capacity from scene appearance toward object dynamics, improving object fidelity though a substantial gap remains. Finally, our conclusions transfer to downstream humanoid modeling. Overall, our results suggest that scaling ego-centric data brings agent modeling close to its limit while leaving its effects on the world far behind, and that closing this gap will depend on how models are trained, not only on how much data they see.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑