发表机构
South China University of Technology; AgiBot(华南理工大学; AgiBot)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
BEE提出干预自适应框架,将人工纠正视为约束证据,通过预测纠正一致性调节策略优化,在真实与仿真任务中超越现有方法,实现高成功率与低干预率。
AI 中文摘要
视觉-语言-动作(VLA)模型能够处理长时程操作任务,然而其成功取决于少数几个精度关键阶段,在这些阶段中毫米级误差会抵消先前所有的进展。在线强化学习(RL)可以精确优化这些动作,但在真实机器人上进行自由探索成本过高,这使得人工纠正变得不可或缺。然而,现有的针对VLA的在线强化学习方法要么无法纳入此类纠正,要么将其混入无差别的监督信号中。实际上,人工纠正并非均匀地带有噪声,而是在某些动作维度上可靠,在其他维度上则变化不定。基于此,我们提出了BEE,一种用于在冻结的VLA上进行真实世界强化学习的干预自适应框架,使策略能够超越专家模仿。我们不是将人工纠正视为要复现的动作,而是将其视为关于约束的证据:一个纠正模型预测人类将如何纠正给定的VLA提议,以及该纠正沿每个动作维度的一致性程度。这种预测的一致性设定了策略优化约束的逐维度紧度。在纠正一致的地方,策略保持接近人类行为;在纠正变化的地方,约束则放宽。我们在三个真实世界操作任务和一个LIBERO-Pro仿真任务上,以匹配的在线数据预算评估了BEE。BEE在每个任务上都取得了最高成功率,平均为91.2%,而RLT为57.5%,DSRL为42.1%,并且在所有真实世界任务上实现了最低的人工干预率。
英文摘要
Vision-language-action (VLA) models handle long-horizon manipulation, yet success hinges on a few precision-critical phases where millimeter-scale errors undo all prior progress. Online reinforcement learning (RL) can optimize exactly these actions, but free exploration is far too costly on real robots, which makes human corrections indispensable. However, existing online RL methods for VLAs either cannot incorporate such corrections or fold them into undifferentiated supervision. Yet human corrections are not uniformly noisy but reliable along some action dimensions and variable along others. Building on this, we introduce BEE, an intervention-adaptive framework for real-world RL on a frozen VLA that lets the policy go BEyond Expert imitation. We formulate human corrections not as actions to reproduce but as evidence about a constraint: a Correction Model predicts how a human would correct a given VLA proposal and how consistent the correction is along each action dimension. This predicted consistency sets the per-dimension tightness of a constraint on policy optimization. Where corrections are consistent the policy stays close to the human, and where they vary, the constraint relaxes. We evaluate BEE on three real-world manipulation tasks and one LIBERO-Pro simulation task at a matched online-data budget. BEE attains the highest success rate on every task, 91.2% on average against 57.5% for RLT and 42.1% for DSRL, and the lowest human intervention rate on all real-world tasks.