面向人在回路在线机器人学习的Max-Q选择性模仿
Max-Q Selective Imitation for Human-in-the-Loop Online Robot Learning
浏览论文内容
中文总结 AI 辅助
本文提出Max-Q选择性模仿方法,结合MC Q-chunk评论家,在真实机器人USB抓取插入等任务中,训练效率显著优于HIL-SERL等基线方法。
中文摘要 AI 辅助
面向真实机器人的人在回路(HIL)在线强化学习必须快速吸收人类干预,同时持续超越人类先验知识。本文针对该场景提出一种基于两个组件的训练方法:其一,MC Q-chunk评论家将 chunk 级动作值回归到回放缓冲器中的蒙特卡洛回报,执行样本平均(行为)策略评估,从而直接为干预轨迹分配信用,而非被当前策略的时间差分(TD)备份稀释;其二,Max-Q选择性模仿通过硬胜者通吃规则,在每个状态下模仿当前策略动作与缓冲器样本中Q值更高的动作,以此更新演员。该规则可在从干预学习与在线策略自我改进之间自动切换:当自主策略更强时,目标与策略分布对齐,减少否则会导致执行时分布偏移的策略-目标样本差距。实际应用中,我们用标准评论家集成均值对候选进行评分以降低比较噪声,无需软化目标或引入分数差距阈值。在包含20个演示的真实USB抓取插入任务中,ACT QChunk-MCBC在30分钟的HIL训练内达到99%的成功率,而HIL-SERL则需要约5小时才能收敛。在Peg Insertion和Square的仿真环境中,ACT/Flow Q-chunk变体在约半小时的有效训练内同样达到≥96%的成功率,在成功率-时间前沿上优于HIL-SERL、EXPO和E2HiL。
英文摘要
Human-in-the-loop (HIL) online reinforcement learning for real robots must absorb human interventions quickly while continuing to improve beyond the human prior. We present a training method for this setting based on two components. First, an \emph{MC Q-chunk} critic regresses chunk-level action values onto Monte Carlo returns from the replay buffer, performing sample-average (behavior) policy evaluation so that intervention trajectories are credited directly rather than diluted by current-policy TD backups. Second, \emph{max-Q selective imitation} updates the actor by imitating, at each state, the higher-$Q$ action between the current policy action and a buffer sample under a hard winner-take-all rule. This rule automatically switches between learning from interventions and on-policy self-improvement: when the autonomous policy is stronger, targets align with the policy distribution, reducing the policy--target-sample gap that otherwise induces execution-time distribution shift. In practice we score candidates with a standard critic ensemble mean to reduce comparison noise, without softening targets or introducing score-gap thresholds. On a real USB pick-and-insertion task with 20 demonstrations, ACT QChunk-MCBC attains 99\% success within 30 minutes of HIL training, whereas HIL-SERL requires about 5 hours to converge. In simulation on Peg Insertion and Square, ACT/Flow Q-chunk variants similarly reach $\ge$96\% success within roughly half an hour of effective training, outperforming HIL-SERL, EXPO, and E2HiL on the success--time frontier.