arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

像世界模型一样思考,像VLA一样行动:将世界模型表征蒸馏进紧凑型机器人策略

Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies

Trung Dao, Sankalp Yamsani, Jaden Park, Joohyung Kim, Yong Jae Lee

arXiv 2609.24682首次发表:更新:

发表机构

University of Wisconsin–Madison; University of Illinois Urbana-Champaign(威斯康星大学麦迪逊分校; 伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

通过将冻结世界模型的内部特征蒸馏到VLA策略中,在不增加推理开销的情况下提升机器人操作的鲁棒性与性能,并在仿真和真实硬件上验证了方法的有效性。

AI 中文摘要

视觉-语言-动作(VLA)模型将观测映射到动作,但缺乏考虑世界如何响应的目标函数,因此其鲁棒性主要受限于数据覆盖范围。世界模型恰好具备这一缺失的目标,并因此具有更好的基础,然而向前滚动未来每次决策需要数秒时间,这使其无法进入控制回路。我们证明这两者可以分离。世界模型对物理场景的了解存在于其内部特征中;生成未来仅仅是产生这些特征的目标,因此可以在舍弃生成机制的同时继承其基础。我们在常规VLA训练中添加了一个特征对齐项:将冻结的世界模型在训练帧上运行一次并缓存,学生模型学习与该缓存保持一致。训练期间不加载教师模型,投影器在训练后被丢弃,部署的策略与未蒸馏的基线完全相同,在消费级RTX 5090上以32毫秒和1.86GB运行,因此所有增益都归因于表征而非增加的容量或测试时计算。一个0.8B的学生模型在LIBERO上达到97.9%,在RoboCasa-GR1人形操作上从48.2%提升至50.5%,且同一目标可迁移至真实硬件,包括单臂和双臂平台。该增益在学生规模、骨干网络、对齐层和教师模型的变化下依然存在,表明这是一种广泛的表征先验,而非两个特定网络之间的脆弱对齐。项目页面:此https URL。

英文摘要

Vision-Language-Action (VLA) models map observations to actions with no objective that accounts for how the world responds, so their robustness is bounded primarily by data coverage. World models carry precisely that missing objective and are better grounded for it, yet rolling the future forward costs seconds per decision and rules them out of the control loop. We show the two can be separated. What a world model knows about physical scenes lives in its internal features; generating the future is merely the objective that produced them, so the grounding can be inherited while the generative machinery is left behind. We add one feature-alignment term to ordinary VLA training: a frozen world model is run over the training frames once and cached, and the student learns to agree with that cache. No teacher is loaded during training, the projector is discarded after it, and the deployed policy is identical to the undistilled baseline, running in 32ms and 1.86GB on a consumer RTX5090, so every gain is attributable to the representation rather than to added capacity or test-time compute. A 0.8B student reaches 97.9% on LIBERO, improves from 48.2% to 50.5% on RoboCasa-GR1 humanoid manipulation, and the same objective carries over to real hardware, on both a single-arm and a bimanual platform. The gain survives changes of student scale, backbone, alignment layer, and teacher, indicating a broad representational prior rather than a fragile alignment between two particular networks. Project page: https://thaw-vla.trung-dt.com/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑