CARF:失败引导流匹配的对比吸引-排斥框架
CARF: Contrastive Attraction-Repulsion of Failure-Guided Flow Matching
浏览论文内容
中文总结 AI 辅助
提出CARF框架,通过进度评分器区分失败轨迹中的渐进与失败关键行为,以流匹配目标吸引策略学习成功行为并排斥失败行为,提升不完美机器人数据利用效率。
中文摘要 AI 辅助
机器人演示收集过程中,除了成功的演示外,常常会产生不完美或失败的轨迹。现有方法通常通过识别那些仍能推进任务完成的片段来利用失败轨迹,但很大程度上忽略了直接导致任务失败的“失败关键行为”。本文认为,这两类片段提供了根本不对称的监督信号:渐进片段应当被模仿,而失败关键片段则应被明确避免。基于这一观察,我们提出了CARF,一种用于从不完美机器人数据中学习的对比吸引-排斥失败引导框架。CARF引入了一个基于进度的重要性评分器,仅通过成功的专家演示及其扰动结果进行训练,以估计每一步对任务完成的贡献,并识别失败轨迹中的信息丰富区域。这些评分指导一个统一的流匹配目标,使策略被吸引向渐进行为,同时排斥失败关键行为,并排除模糊片段。这使得能够更全面地利用不完美数据,并避免来自模糊失败片段的不可靠监督。在仿真和真实世界中的大量实验表明,在多种失败场景下,CARF相比竞争基线取得了一致的改进,消融研究进一步验证了所提出的评分和吸引-排斥机制的有效性。我们的网站是此https URL。
英文摘要
Robot demonstration collection often produces imperfect or failed trajectories in addition to successful demonstrations. Existing methods typically exploit failed trajectories by identifying segments that still make progress toward task completion, but largely overlook \textit{failure-critical behaviors} that directly lead to task failure. Here we argue that these two types of segments provide fundamentally asymmetric supervision: progressive segments should be imitated, whereas failure-critical segments should be explicitly avoided. Based on this observation, we propose CARF, a Contrastive Attraction-Repulsion of Failure-guided framework for learning from imperfect robot data. CARF introduces a progress-based importance scorer, trained solely on successful expert demonstrations and its perturbation results, to estimate step-wise contributions toward task completion and identify informative regions in failed trajectories. These scores guide a unified flow-matching objective that attracts the policy toward progressive behaviors and repels it from failure-critical ones, while excluding ambiguous segments. This enables more comprehensive utilization of imperfect data and avoids unreliable supervision from ambiguous failure segments. Extensive experiments in simulation and the real world demonstrate consistent improvements over competing baselines across diverse failure scenarios, with ablations further validating the effectiveness of the proposed scoring and attraction-repulsion mechanisms. Our website is https://zhao-sq.github.io/carf/#.