发表机构
University of Science and Technology of China; Beihang University; Zhongguancun Academy; Hefei SpinX Technology(中国科学技术大学; 北京航空航天大学; 中关村学院; 合肥星旋科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉-语言-动作(VLA)模型在线RL的策略漂移与效率瓶颈,提出含ACoB算法及ACoB-Stream架构的VLA-Precision框架,在9项高精度化学任务上实现高成功率与效率提升。
AI 中文摘要
预训练的视觉-语言-动作(VLA)模型可实现广泛的操作任务,但在要求精度和可重复性的任务中仍不可靠。将真实世界在线强化学习(RL)应用于VLA的后训练,可实现仅通过演示之外的自主试错改进,但存在两个瓶颈:1)不可靠的价值信号会导致策略漂移;2)大型VLA的开销会限制吞吐量和样本效率。为应对这些挑战,我们提出VLA-Precision,这是一种高效的真实世界在线RL框架,包含非对称自举引导(ACoB)算法和ACoB-Stream架构。具体而言,ACoB在时间尺度上建立非对称自举引导:早期干预引导的行为学习可快速提升策略性能,同时提高在线经验质量。随着自主经验的积累,全局回报传播和局部偏好排序会逐步校准价值估计,为参考正则化的策略改进提供相对动作优势,同时抑制漂移。为使ACoB能应用于大型VLA,我们开发了ACoB-Stream,这是一种闭环经验-策略架构,以不变态解耦和按需流处理为设计原则,可实现吞吐量和计算效率提升达10.9倍。在四个类别、四个机器人 embodiment的九个高精度化学任务上进行的广泛评估显示,VLA-Precision在45.8分钟/任务内达到98.3%的平均成功率,其27.6秒的回合运行速度分别是VLA和RL基准的1.2倍和1.8倍。资源可在此https URL获取。
英文摘要
Pretrained vision-language-action (VLA) models enable broad manipulation but remain unreliable in tasks demanding precision and repeatability. Applying real-world online reinforcement learning (RL) to VLA post-training enables autonomous trial-and-error improvement beyond demonstrations alone, but exposes two bottlenecks: 1) unreliable value signals can induce policy drift; 2) large-VLA overhead constrains throughput and sample efficiency. To address these challenges, we present VLA-Precision, an efficient real-world online RL framework featuring the Asymmetric Co-Bootstrapping (ACoB) algorithm and the ACoB-Stream architecture. Specifically, ACoB establishes asymmetric co-bootstrapping across timescales: early intervention-guided behavioral learning rapidly improves policy performance while enhancing online experience quality. As autonomous experience accumulates, global return propagation and local preference ranking progressively calibrate value estimates, yielding relative action advantages for reference-regularized policy improvement while suppressing drift. To enable ACoB on large VLAs, we develop ACoB-Stream, a closed-loop experience--policy architecture that establishes invariant-state decoupling and on-demand streaming as design principles, delivering up to 10.9$\times$ improvements in throughput and computational efficiency. Extensive evaluations on nine high-precision chemistry tasks across four categories and four robot embodiments show that VLA-Precision achieves 98.3\% mean success rate in 45.8 min/task, with 27.6 s episodes running at 1.2$\times$ and 1.8$\times$ the speeds of VLA and RL baselines. Resources are available at https://vla-precision.github.io.
Comments17 pages, 14 figures