arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

剖析优势引导的视觉-语言-动作策略后训练

Dissecting Advantage-Guided Post-Training for Vision-Language-Action Policies

Jiahang Cao, Hanye Zhao, Hang Lai, Shenyu Zhang, Xiaoshen Han, Xinghang Li, Futeng Liu, Wanli Peng, Heyun Wang, Yunhong Wang, Jason Li, Yong Yu, Weinan Zhang

arXiv 2609.28161首次发表:更新:

发表机构

School of Computer Science, Shanghai Jiao Tong University; Xiaomi Robotics(上海交通大学计算机科学学院; 小米机器人)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过受控实证研究剖析优势引导的VLA策略后训练,提出模块化方案,在真实双臂任务上显著提升性能。

AI 中文摘要

优势引导的强化学习提供了一种利用有限机器人数据对视觉-语言-动作(VLA)策略进行后训练的实用方法。然而,其性能依赖于多个耦合的选择,包括如何构建、校准评论家派生的优势以及如何将其用于策略训练。现有的方案通常将这些选择合并到一个端到端的流程中,使得难以识别它们各自的影响。在这项工作中,我们通过一项受控的实证研究来剖析优势引导的VLA后训练,该研究分离了这些设计选择,同时考虑了它们不同的估计目标。我们开发了分阶段的离线评估方法,以高效筛选替代选择,而无需对每种可能的组合进行大量的真实机器人策略评估。分阶段评估确定了一个模块化方案,该方案结合了时间差分优势构建、分组校准和连续优势加权。在四个真实世界的双臂任务中,所得方案相对于SFT初始化将平均任务进度和成功率分别提高了0.42和0.63。此外,所提出的评估诊断显示出与下游真实世界性能的整体一致性,支持其在实践中用于解释经验结果和选择优势引导的后训练设计。

英文摘要

Advantage-guided reinforcement learning provides a practical way to post-train vision-language-action (VLA) policies using limited robot data. However, its performance depends on several coupled choices, including how critic-derived advantages are constructed, calibrated, and used for policy training. Existing recipes often combine these choices into a single end-to-end procedure, making their individual effects difficult to identify. In this work, we dissect advantage-guided VLA post-training through a controlled empirical study that separates these design choices while accounting for their distinct estimands. We develop stage-specific offline evaluation methods to screen alternative choices efficiently, without requiring extensive real-robot policy evaluations for every possible combination. The staged evaluation identifies a modular recipe that combines temporal-difference advantage construction, group-wise calibration, and continuous advantage weighting. Across four real-world bimanual tasks, the resulting recipe improves mean task progress and success over the SFT initialization by 0.42 and 0.63, respectively. Moreover, the proposed evaluation diagnostics show an overall alignment with downstream real-world performance, supporting their use for interpreting empirical outcomes and selecting advantage-guided post-training designs in practice.

Comments9 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑