ForeTime-VLA:面向传送带操作的世界动作模型的因果未来令牌蒸馏
ForeTime-VLA: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation
浏览论文内容
中文总结 AI 辅助
该研究提出ForeTime-VLA策略,通过从冻结Fast-WAM教师模型蒸馏因果未来令牌,在传送带操控任务上降低了MAE和L2误差,提升了抓取成功率,验证了该方法的有效性。
中文摘要 AI 辅助
操控移动物体需要策略预判接触事件,但视觉-语言-动作(VLA)策略通常仅基于当前观测进行微调。世界动作模型(WAM)可学习预测动力学,但部署时运行视频规模教师模型或显式想象未来帧成本高昂。我们提出ForeTime-VLA,这是一种密集pi0.5策略,从冻结的Fast-WAM衍生教师模型中蒸馏出具有未来感知、动作等价的表示,同时在推理时保持因果性。离线阶段,将当前和未来视频潜变量压缩为白化的64维目标;在线阶段,使用8帧历史编码器预测该目标,以及操作阶段和归一化的转换时间。4个未来令牌和1个阶段令牌条件化VLM前缀,预测的未来和转换视界条件化动作专家。训练保留原始流匹配动作目标,并添加余弦、关系几何、阶段、转换时间和动作等价性目标。在去重的传送带数据集上,我们在每个拆分的768个匹配窗口上比较4万步检查点:测试MAE从0.134119降至0.130593(降幅2.63%;配对自助法95%置信区间:0.82-4.48%的提升),测试L2降幅3.02%,延迟成本为2.46-2.93%。在定量真实机器人评估中,ForeTime-VLA实现81.1%的静止和58.9%的慢移抓取成功率,分别比次优参考高出12.2和22.2个百分点。在三种传送带速度下,其完成44/90次抓取,而pi0.5为23/90次,其中高速时为11/30次对比2/30次。离线姿态提升与真实机器人接触位姿失败减少之间的一致性,支持因果未来令牌蒸馏是改进动态操控、无需部署世界模型教师的有效方法。
英文摘要
Manipulating moving objects requires a policy to anticipate contact events, yet vision-language-action (VLA) policies are commonly fine-tuned from the current observation alone. World action models (WAMs) learn predictive dynamics, but running a video-scale teacher or explicitly imagining future frames at deployment is costly. We introduce ForeTime-VLA, a dense pi0.5 policy that distills a future-aware, action-equivalent representation from a frozen Fast-WAM-derived teacher while remaining causal at inference. Offline, current and future video latents are compressed into a whitened 64-D target. Online, an eight-frame history encoder predicts this target together with manipulation phase and normalized time-to-transition. Four future tokens and one phase token condition the VLM prefix, while the predicted future and transition horizon condition the action expert. Training retains the original flow-matching action target and adds cosine, relational geometry, phase, time-to-transition, and action-equivalence objectives. On a deduplicated conveyor-belt dataset, we compare 40k-step checkpoints on 768 matched windows per split. Test MAE decreases from 0.134119 to 0.130593 (2.63%; paired-bootstrap 95% CI: 0.82-4.48% improvement), and test L2 decreases by 3.02%, at a 2.46-2.93% latency cost. In quantitative real-robot evaluation, ForeTime-VLA achieves 81.1% stationary and 58.9% slow-moving grasp success, exceeding the next-best reference by 12.2 and 22.2 percentage points, respectively. Across three belt speeds, it completes 44/90 grasps versus 23/90 for pi0.5, including 11/30 versus 2/30 at fast speed. The agreement between offline orientation gains and reduced real-robot contact-pose failures supports causal future-token distillation as an effective way to improve dynamic manipulation without deploying the world-model teacher.
发表机构
- Tsinghua University(清华大学)
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
- Harbin Institute of Technology(哈尔滨工业大学)
- Hangzhou Yunshenchu Technology Co., Ltd. (DEEP Robotics)(杭州云深处科技有限公司(深地机器人))
机构由 AI 辅助整理,请以论文原文为准。