arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TOUR:用于离线强化学习的轨迹级遗忘基准测试

TOUR: A Trajectory-Level Unlearning Benchmark for Offline Reinforcement Learning

Chaofan Pan, Lingfei Ren, Xiangyu Jiang, Yanhua Li, Xuemei Cao, Xiangkun Wang, Hao Yu, Wei Wei, Xin Yang

arXiv 2607.21111首次发表:更新:

发表机构

Southwestern University of Finance and Economics; Shanxi University(西南财经大学; 山西大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对离线强化学习中轨迹级删除难评估的问题,提出TOUR基准测试,结合多种技术,通过实验表明常见删除基线有环境依赖隐私 - 效用行为,单一分数审计下结论不稳定,其取决于多种因素。

AI 中文摘要

离线强化学习(RL)智能体基于固定行为轨迹进行训练,这使得在训练后必须删除选定数据时,轨迹级删除变得很重要。评估这种删除很困难,因为较低的成员分数可能反映轨迹删除、对另一种攻击可见的残留记忆或破坏有用行为的策略崩溃。我们引入了离线RL中的轨迹级记忆与遗忘(TOUR)基准测试,它结合了轨迹级分区、匹配的非成员控制、重新训练参考、保留性能锚点和多攻击隐私审计。在D4RL运动实验和探索性的蚂蚁迷宫扩展实验中,TOUR表明常见的删除基线具有依赖环境的隐私 - 效用行为。重新训练和微调通常比统一的遗传算法+重新拟合提供更强的保留效用参考,而TrajDeleter仍然是一个有用的比较器,但在相同审计下并非始终更强。参考模型、阈值、偏差、等价性、动作误差、基于表示和查询限制的攻击进一步表明,基于单一似然的成员分数可能高估删除质量。因此,在评估设置中,关于离线RL遗忘的结论在单分数审计下不稳定。它们取决于匹配的非成员构建、重新训练相对校准、攻击家族、保留效用以及诊断架构或组件级证据的明确范围。

英文摘要

Offline Reinforcement Learning (RL) agents are trained on fixed behavioral trajectories, which makes trajectory-level deletion important when selected data must be removed after training. Evaluating such deletion is difficult because a lower membership score can reflect trajectory removal, residual memorization visible to another attack, or policy collapse that destroys useful behavior. We introduce Trajectory-level memOrization and Unlearning in offline RL (TOUR), a benchmark that combines trajectory-level partitioning, matched non-member controls, retraining references, retained-performance anchors, and multi-attack privacy auditing. Across D4RL locomotion experiments and an exploratory AntMaze extension, TOUR shows that common deletion baselines have environment-dependent privacy-utility behavior. Retraining and fine-tuning often provide stronger retained-utility references than uniform GA+Refit, while TrajDeleter remains a useful comparator but is not uniformly stronger under the same audit. Reference-model, threshold, deviation, equivalence, action-error, representation-based, and query-limited attacks further show that a single likelihood-based membership score can overstate deletion quality. In the evaluated settings, conclusions about offline RL unlearning are therefore not stable under single-score auditing. They depend on matched non-member construction, retraining-relative calibration, attack family, retained utility, and explicit scope for diagnostic architecture or component-level evidence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑