arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

REDIRECT:机器人不良习惯的1%修复方案

REDIRECT: A 1% Fix for Bad Robot Habits

Yu Zhang, Jiazhuo Li, Yancong Wei, Kangkang Dong, Xiaojun Zhu, Houde Liu

arXiv 2610.03997首次发表:更新:

发表机构

Tsinghua University; University of Michigan(清华大学; 密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

REDIRECT利用仅1%的更新预算,通过定位并修复局部行为差异,将机器人成功率从68.2%提升至90.3%,高效修复不良习惯。

AI 中文摘要

机器人可能会从原本有用的遥操作数据中的少数缺陷时刻习得不良习惯。在混合质量的机器人数据中,正常演示与问题演示共享大部分任务行为,仅在局部动作延续上存在差异。完全重新训练成本高昂,而仅对干净数据进行微调在较小的更新预算下恢复效果有限。我们探究是否可以用全训练样本反向传播预算的仅1%来修复这种局部差异。仅使用基于回合的保留/问题标签,REDIRECT定位分支,为周围原始观测分配一致的保留延续,并锚定共享行为,无需帧级标注或额外交互。在三个随机ManiSkill任务和三个随机种子上,REDIRECT将平均成功率从68.2%提升至90.3%,恢复了干净重训练差距的88.8%,而匹配计算量的微调仅恢复22.1%。在PiPER机械臂上,在杯子插入和毛巾折叠任务中,它恢复了干净数据训练差距的86.7%-92.9%。因此,局部机器人习惯可以通过将更新预算用于行为差异而非重新学习共享行为来修复。

英文摘要

Robots can acquire bad habits from a few defective moments in otherwise useful teleoperation. In mixed-quality robot data, normal and problematic demonstrations share most task behavior and differ only at a local action continuation. Full retraining is costly, while fine-tuning on clean data alone offers limited recovery under a small update budget. We ask whether the local difference can instead be fixed with only 1% of the full-training sample-backward budget. Using only episode-level retained/problematic labels, REDIRECT localizes the branch, assigns a coherent retained continuation to the original observations around it, and anchors shared behavior, without frame-level annotations or additional interaction. Across three randomized ManiSkill tasks and three seeds, REDIRECT raises the mean success rate from 68.2% to 90.3%, recovering 88.8% of the clean-retraining gap versus 22.1% for matched-compute fine-tuning. On a PiPER arm, it recovers 86.7-92.9% of the clean-only gap across cup insertion and towel folding. Local robot habits can therefore be repaired by spending the update budget on the behavioral difference rather than relearning shared behavior.

Comments8 pages, 5 figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑