发表机构
Southwestern University of Finance and Economics; Shanxi University(西南财经大学; 山西大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对离线强化学习中轨迹级删除难评估的问题,提出TOUR基准测试,结合多种技术,通过实验表明常见删除基线有环境依赖隐私 - 效用行为,单一分数审计下结论不稳定,其取决于多种因素。
AI 中文摘要
离线强化学习(RL)智能体基于固定行为轨迹进行训练,这使得在训练后必须删除选定数据时,轨迹级删除变得很重要。评估这种删除很困难,因为较低的成员分数可能反映轨迹删除、对另一种攻击可见的残留记忆或破坏有用行为的策略崩溃。我们引入了离线RL中的轨迹级记忆与遗忘(TOUR)基准测试,它结合了轨迹级分区、匹配的非成员控制、重新训练参考、保留性能锚点和多攻击隐私审计。在D4RL运动实验和探索性的蚂蚁迷宫扩展实验中,TOUR表明常见的删除基线具有依赖环境的隐私 - 效用行为。重新训练和微调通常比统一的遗传算法+重新拟合提供更强的保留效用参考,而TrajDeleter仍然是一个有用的比较器,但在相同审计下并非始终更强。参考模型、阈值、偏差、等价性、动作误差、基于表示和查询限制的攻击进一步表明,基于单一似然的成员分数可能高估删除质量。因此,在评估设置中,关于离线RL遗忘的结论在单分数审计下不稳定。它们取决于匹配的非成员构建、重新训练相对校准、攻击家族、保留效用以及诊断架构或组件级证据的明确范围。
英文摘要
Offline Reinforcement Learning (RL) agents are trained on fixed behavioral trajectories, which makes trajectory-level deletion important when selected data must be removed after training. Evaluating such deletion is difficult because a lower membership score can reflect trajectory removal, residual memorization visible to another attack, or policy collapse that destroys useful behavior. We introduce Trajectory-level memOrization and Unlearning in offline RL (TOUR), a benchmark that combines trajectory-level partitioning, matched non-member controls, retraining references, retained-performance anchors, and multi-attack privacy auditing. Across D4RL locomotion experiments and an exploratory AntMaze extension, TOUR shows that common deletion baselines have environment-dependent privacy-utility behavior. Retraining and fine-tuning often provide stronger retained-utility references than uniform GA+Refit, while TrajDeleter remains a useful comparator but is not uniformly stronger under the same audit. Reference-model, threshold, deviation, equivalence, action-error, representation-based, and query-limited attacks further show that a single likelihood-based membership score can overstate deletion quality. In the evaluated settings, conclusions about offline RL unlearning are therefore not stable under single-score auditing. They depend on matched non-member construction, retraining-relative calibration, attack family, retained utility, and explicit scope for diagnostic architecture or component-level evidence.