arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.03681cs.RO

WISE:用于视觉-语言-动作模型高效后训练的世界模型引导式想象调度

WISE: World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models

Chenhao Zhang, Hanyu Zhao, Hang Cheng, Tengfei Pan, Long Zeng

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出WISE框架,通过世界模型引导的想象调度优化VLA模型后训练,在多操作任务上提升性能,减少约80%GPU计算时间,增强真实世界下的鲁棒性与泛化性。

中文摘要 AI 辅助

视觉-语言-动作(VLA)策略的后训练通常依赖于使用代价高昂的专家演示进行监督微调,或依赖于在真实世界中进行代价高昂且可能不稳定的探索的强化学习。世界模型通过想象未来来评估候选行为,提供了一种有前景的替代方案,但有效的后训练不仅需要准确的预测:想象必须在有用的地方调度,限制在可靠的范围内,并转化为可信的策略监督。在机器人操作中,想象的价值在不同执行阶段差异很大,而扩展的回滚会累积预测误差并引入不可靠的学习信号。我们提出了WISE(World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models),这是一个在策略优化过程中协调何时以及如何使用世界模型想象的统一框架。WISE在与交互相关的状态下选择性调用想象,执行有界的多视角回滚,使用进度和完成信号评估候选未来,并利用它们的相对结果来优化从真实交互上下文生成的动作。对π₀和π₀.⁵的大量实验表明,在各种操作任务上取得了一致的改进,同时与完全想象相比,减少了约80%的GPU计算时间。真实世界评估进一步显示,在各种真实世界分布偏移下,鲁棒性和泛化能力有显著提升。

英文摘要

Post-training VLA policies typically rely on supervised fine-tuning with costly expert demonstrations or reinforcement learning with expensive and potentially unstable real-world exploration. World models offer a promising alternative by evaluating candidate behaviors through imagined futures, yet effective post-training requires more than accurate prediction: imagination must be scheduled where it is useful, bounded within reliable horizons, and translated into trustworthy policy supervision. In robotic manipulation, the value of imagination varies substantially across execution stages, while extended rollouts can accumulate prediction errors and introduce unreliable learning signals. We introduce WISE (World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models), a unified framework that coordinates when and how world-model imagination is used during policy refinement. WISE selectively invokes imagination at interaction-relevant states, performs bounded multi-view rollouts, evaluates candidate futures using progress and completion signals, and uses their relative outcomes to refine actions generated from real interaction contexts. Extensive experiments with both $π_0$ and $π_{0.5}$ demonstrate consistent improvements across diverse manipulation tasks while reducing GPU computation time by approximately 80% compared with full imagination. Real-world evaluations further show substantial gains in robustness and generalization under diverse real-world distribution shifts.

发表机构

  • Tsinghua University(清华大学)
  • Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院(BAAI))

机构由 AI 辅助整理,请以论文原文为准。

↑