arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在线交互对于恢复是否必要?一种通过扰动的极简鲁棒规划方法

Is Online Interaction Necessary for Recovery? A Minimalist Approach to Robust Planning via Perturbation

Bumgeun Park, Donghwan Lee

arXiv 2609.33049首次发表:更新:

AI 中文总结

针对行为克隆在闭环执行中易受协变量偏移影响的问题,提出扰动增强恢复监督(PARS),仅用现有演示通过扰动状态并锚定轨迹后半段来提供恢复监督,无需额外交互或专家,在51个RLBench任务上验证了鲁棒性提升。

AI 中文摘要

行为克隆(BC)在闭环执行过程中容易受到协变量偏移的影响,此时小的预测或执行误差可能将机器人驱向演示数据覆盖不佳的状态。我们聚焦于动作序列规划,其中策略预测有限时域的动作序列作为机器人执行的参考轨迹。现有的协变量偏移处理方法通常依赖收集额外的纠正性演示,这需要进一步的环境交互以及专家或参考策略的访问。我们提出扰动增强恢复监督(PARS),一种仅利用现有演示来改善恢复能力的简单训练方法。PARS扰动机器人的本体感受状态,并仅将预测轨迹的后半部分锚定到原始演示,使前半部分自由生成纠正运动,并在规划时域内恢复至演示行为。与传统的输入噪声增强(在整个预测时域内保留原始监督目标)不同,PARS明确为从扰动状态恢复提供轨迹级监督。PARS既不需要额外的环境交互,也不需要专家查询,并且只需对行为克隆目标进行少量修改即可实例化到不同的动作序列策略类别中。在51个RLBench操作任务上,使用基于流、基于变换器和基于扩散的策略进行的实验表明,PARS提高了闭环执行中对协变量偏移的鲁棒性。

英文摘要

Behavior cloning (BC) is vulnerable to covariate shift during closed-loop execution, where small prediction or execution errors can drive the robot toward states poorly covered by the demonstration data. We focus on action-sequence planning, where a policy predicts a finite-horizon sequence of actions as a reference trajectory for robot execution. Existing approaches to covariate shift often rely on collecting additional corrective demonstrations, requiring further environment interaction and access to an expert or reference policy. We propose Perturbation-Augmented Recovery Supervision (PARS), a simple training approach for improving recovery using only existing demonstrations. PARS perturbs the robot's proprioceptive state and anchors only the later portion of the predicted trajectory to the original demonstration, leaving the earlier portion free to generate corrective motion and recover toward the demonstrated behavior within the planning horizon. Unlike conventional input-noise augmentation, which preserves the original supervision target over the entire prediction horizon, PARS explicitly provides trajectory-level supervision for recovery from perturbed states. PARS requires neither additional environment interaction nor expert queries and can be instantiated across different action-sequence policy classes with only minor modifications to their BC objectives. Experiments on 51 RLBench manipulation tasks with flow-based, transformer-based, and diffusion-based policies demonstrate that PARS improves robustness to covariate shift during closed-loop execution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑