arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22888cs.RO

稳定高效的现实世界在线VLA后训练:基于异步重放锚定策略改进

Stable and Efficient Real-World Online VLA Post-Training via Asynchronous Replay-Anchored Policy Improvement

Jiarui Yang, Jiajin Zhang, Bin Zhu, Jingjing Chen, Yu-Gang Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

提出异步重放锚定策略改进(RAPolicy)框架,通过并行数据采集与学习、基于重放行为的评论家和演员更新,实现VLA模型高效在线后训练,在真实世界多任务中成功率从52%提升至88%。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型的在线后训练需要高效利用机器人交互,并从持续收集的经验中实现可靠的政策改进。我们提出了异步重放锚定策略改进(RAPolicy)框架,该框架并行执行数据采集和学习,同时将评论家(critic)和演员(actor)的更新锚定在重放的行为上。评论家从记录的动作中学习块级价值,并构建贝尔曼目标而无需预测下一动作,从而减少计算量并降低对重放覆盖范围之外的动作价值估计的依赖。一步流演员复用数据采集期间存储的初始噪声,并通过优势加权条件似然进行学习,直接监督用于执行的行动映射。我们在现实世界中评估了RAPolicy,涵盖四个单任务场景和一个联合五任务场景,在线训练预算仅为1至2小时。从每个任务仅用10个演示微调的初始策略出发,RAPolicy快速适应新的单一任务,平均成功率达到86.3%。在联合多任务场景中,RAPolicy将整体成功率从52%提升至88%,同时保持已有可靠任务的性能并增强较弱的能力。总体而言,RAPolicy在总任务成功率上大幅优于基线,同时需要更少的人工干预,展示了稳定的政策改进和高效的在线训练。项目页面:此https URL。

英文摘要

Online post-training of vision-language-action (VLA) models requires efficient use of robot interaction and reliable policy improvement from continually collected experience. We propose asynchronous Replay-Anchored Policy improvement (RAPolicy), a framework that performs rollout and learning concurrently while grounding both critic and actor updates in replayed behavior. The critic learns chunk-level values from recorded actions and constructs Bellman targets without predicting next actions, reducing computation and dependence on action-value estimates outside replay coverage. The one-step flow actor reuses the initial noise stored during rollout and learns through advantage-weighted conditional likelihood, directly supervising the action mapping used for execution. We evaluate RAPolicy across four single-task settings and one joint five-task setting in the real world, with online training budgets of only 1--2 hours. Starting from policies fine-tuned on just 10 demonstrations per task, RAPolicy rapidly adapts to new single tasks and achieves an average 86.3% success rate. In the joint multi-task setting, RAPolicy improves overall success rate from 52% to 88% while preserving performance on already reliable tasks and improving weaker capabilities. Overall, RAPolicy substantially outperforms the baselines in aggregate task success while requiring fewer human interventions, demonstrating stable policy improvement and high online training efficiency. Project page: https://flyfaerss.github.io/RAPolicy.

发表机构

  • Fudan University(复旦大学)
  • Shanghai Key Laboratory of Multimodal Embodied AI(上海多模态具身智能重点实验室)
  • Singapore Management University(新加坡管理大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑