AI 中文总结
该研究提出CounterAlign方法,通过指令重标记从专家演示合成反事实元组,结合对抗判别器训练学习奖励模型,提升VLA模型在LIBERO-PRO基准及TX-G2机器人实验中的鲁棒性与性能。
AI 中文摘要
视觉-语言-动作(VLA)模型通常采用行为克隆(BC)在专家演示上进行训练。然而,BC仅为专家动作提供正监督,没有明确的负监督来指示哪些动作与指令不一致或不合适。强化学习(RL)可以提供此类校正信号,但通常依赖外部指定的奖励或精心策划的非专家数据,这两者在机器人技术中都难以获取。我们表明,用于VLA模型的离线RL无需依赖精心策划的非专家轨迹:仅通过指令重新标记,成功的专家演示即可转化为密集的校正监督。具体而言,通过将专家动作与不匹配的替代指令配对,我们从数据集中合成反事实的指令-观测-动作元组,并将其与对抗判别器训练相结合,以学习用于离线RL的基于指令的奖励模型,无需收集额外的rollout或注释。在以鲁棒性为核心的LIBERO-PRO基准上,我们的方法在物体位置和任务扰动的鲁棒性方面优于强大的最先进基线。在TX-G2(与AGIBot G2兼容)的真实机器人实验中,它也优于竞争基线。更广泛地说,我们的结果表明,在数据受限的VLA学习中,从每个演示中提取更密集的监督可以补充额外数据的收集。
英文摘要
Vision-Language-Action (VLA) models are typically trained with behavior cloning (BC) on expert demonstrations. However, BC provides only positive supervision for expert actions, without explicit negative supervision indicating which actions are instruction-inconsistent or otherwise inappropriate. Reinforcement learning (RL) can provide such corrective signals, but often relies on externally specified rewards or curated non-expert data, both of which are costly to obtain in robotics. We show that offline RL for VLA models need not rely on curated non-expert trajectories: successful expert demonstrations alone can be transformed into dense corrective supervision through instruction relabeling. Specifically, by pairing expert actions with mismatched alternative instructions, we synthesize counterfactual instruction-observation-action tuples from the dataset and combine them with adversarial discriminator training to learn an instruction-grounded reward model for offline RL, without collecting additional rollouts or annotations. On the robustness-focused LIBERO-PRO benchmark, our method improves robustness to object position and task perturbations over a strong state-of-the-art baseline. It also outperforms competitive baselines in real-robot experiments on the TX-G2 (compatible with AGIBot G2). More broadly, our results suggest that, for data-constrained VLA learning, extracting denser supervision from each demonstration can complement collecting additional data.
CommentsProject page: https://counteralign.airoa.io