机器人模仿学习中演示数据遗忘的重新思考
Rethinking Demonstration Unlearning in Imitation Learning for Robotics
浏览论文内容
中文总结 AI 辅助
本研究针对机器人模仿学习的演示数据遗忘问题,提出双维度联合审计方法,在ACT机械臂实验中实现了20次试验18次成功的盲测恢复。
中文摘要 AI 辅助
机器人模仿学习依赖人类演示数据,其中部分演示数据后续可能被要求移除。不使用这些数据重新训练是自然的参考方案,但重新训练的成本随策略和数据集规模增长,因此催生了可编辑已训练策略的低成本操作。从机器遗忘继承而来的指标(如遗忘损失或单次成员推理攻击)无法确定编辑操作从闭环运行的策略中移除了什么。为此,我们引入了经重新训练校准的审计方法,从两个维度评估演示数据遗忘:行为维度,即编辑后的策略是否与未使用被移除演示数据重新训练的策略表现一致;证据维度,即审计者是否仍能检测到该策略曾基于这些演示数据训练。行为维度通过匹配状态下编辑后策略与重新训练策略的动作差异进行衡量,由独立重新训练得到的基线值校准,处于基线水平的策略与重新训练策略的接近程度,等同于不同重新训练策略之间的接近程度。证据维度对重新训练空模型执行逐演示数据的成员推理攻击,同时报告攻击的排名和绝对成员损失水平,因为仅排名会接受将成员损失膨胀至超过空模型的操作。随后,我们采用共形检验将两个维度结合为联合重新训练一致性假设,通过足够多的独立重新训练(以常规显著性水平可拒绝)进行验证。在三类真实机器人策略和两个仿真套件的五项预注册条件下,某一检查点的两个维度出现双向分离,即编辑操作可在修复任务行为的同时保留证据,或在减少证据的同时使行为偏离重新训练。在ACT机械臂上,重定向编辑将盲测的机器人成功率恢复至20次试验中的18次。
英文摘要
Imitation learning for robotics depends on human demonstrations, some of which people may later ask to remove. Retraining without them is the natural reference, but its cost grows with policy and dataset scale, motivating cheaper operators that edit a trained policy. Metrics inherited from machine unlearning, such as forgetting loss or a single membership attack, do not establish what an edit removed from a policy acting in closed loop. We therefore introduce a retrain-calibrated audit that reads demonstration unlearning along two axes: behavior, whether the edited policy acts like one retrained without the removed demonstrations, and evidence, whether an auditor can still detect it was trained on them. The behavior axis measures action divergence to that retrain at matched states, calibrated by a floor built from independent retrains, so a policy at the floor is as close to a retrain as retrains are to each other. The evidence axis applies a per-demonstration membership attack against a retrain null, reporting both its rank and its absolute member-loss level, since rank alone accepts operators that inflate member losses past the null. A conformal test then combines both axes into one hypothesis of joint retrain consistency, against a fleet of independent retrains large enough to reject at conventional significance. Across five preregistered conditions on three real-robot policy classes and two simulation suites, the axes dissociate in both directions on one checkpoint, as an edit may repair task behavior while leaving evidence unchanged, or reduce evidence while moving behavior away from retraining. On the ACT arm, a redirect edit restores blind-scored robot success to 18 of 20 trials.
发表机构
- University of Michigan(密歇根大学)
- Zhejiang University(浙江大学)
- Wuhan University of Technology(武汉理工大学)
机构由 AI 辅助整理,请以论文原文为准。