不仅仅是更多演示:用于数据高效的鲁棒机器人模仿的反事实动作敏感性覆盖
It's Not Just More Demos: Counterfactual Action Sensitivity Coverage for Data-Efficient Robust Robot Imitation
浏览论文内容
中文总结 AI 辅助
针对视觉运动模仿策略对干扰项脆弱的问题,提出CFNBC框架,通过测量动作漂移选择紧凑修复集,在机器人任务中20-30个候选的修复性能优于随机选择。
中文摘要 AI 辅助
视觉运动模仿学习在操纵任务中已取得成功,但训练出的策略对视觉“干扰项”仍很脆弱,即使是光照、干扰或颜色变化等微小的任务保留型变化,也会导致策略性能严重下降。虽然增加数据多样性可提升鲁棒性,但对于特定训练策略而言,哪些额外演示具有信息价值尚不明确。我们提出反事实干扰行为克隆(Counterfactual Nuisance Behaviour Cloning, CFNBC),这是一种用于针对性鲁棒性修复的离线数据选择框架。从基于“干净”演示训练的名义策略出发,CFNBC生成保留专家动作的配对干净与干扰观测,随后测量“动作漂移”——即策略在本不应改变期望行为的干扰下预测动作的变化。这提供了一种策略特异性敏感性信号,用于从更大的候选池中选择紧凑、响应多样化的修复集,无需部署成功标签或在线策略执行。我们在MuJoCo双机械手立方体搬运任务和SimplerEnv立方体堆叠任务中表明,动作漂移与干扰导致的失败相关,且仅使用20-30个选定候选的响应引导修复,其性能显著优于匹配预算的随机选择,同时接近大得多的随机修复预算的性能。这些结果支持鲁棒性修复的数据中心观点:最有用的数据不一定是数量最多、视觉上最多样或明显最难的,而是覆盖当前策略脆弱响应模式的示例。
英文摘要
Visuomotor imitation learning has demonstrated success for manipulation tasks. However, the trained policies remain brittle to visual `nuisances', with even minor task-preserving variations such as lighting, distractions or changes in colour result in heavy degradation of the trained policy's performance. While increasing data diversity can improve robustness, it is unclear which additional demonstrations are informative for a particular trained policy. We propose Counterfactual Nuisance Behaviour Cloning (CFNBC), an offline data-selection framework for targeted robustness repair. Starting from a nominal policy trained on `clean' demonstrations, CFNBC generates paired clean and nuisance observations that preserve the expert action, then measures \emph{action drift}: the change in the policy's predicted action under a nuisance that should not alter the desired behaviour. This provides a policy-specific sensitivity signal for selecting a compact, response-diverse repair set from a larger candidate pool, without requiring rollout success labels or online policy execution. We show in MuJoCo bimanual cube transfer and SimplerEnv cube stacking that action drift correlates with nuisance-induced failure, and that response-guided repair with only $20$--$30$ selected candidates substantially outperforms matched-budget random selection while approaching the performance of much larger random repair budgets. These results support a data-centric view of robustness repair: the most useful data are not necessarily the most numerous, visually diverse, or obviously difficult, but the examples that cover fragile response modes of the current policy.
发表机构
- CSIRO Robotics, CSIRO, Australia(澳大利亚联邦科学与工业研究组织机器人学部,澳大利亚联邦科学与工业研究组织)
机构由 AI 辅助整理,请以论文原文为准。