UBA-ORL:离线强化学习中的反学习激活后门攻击
UBA-ORL: Unlearning-Activated Backdoor Attacks on Offline Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
针对离线强化学习提出首个反学习激活后门攻击UBA-ORL,利用双样本机制在数据删除后重新激活后门,揭示合规删除流程中的安全漏洞。
中文摘要 AI 辅助
离线强化学习(offline RL)能够从预先收集的静态数据集中学习策略,而无需在线探索,并且不仅越来越多地部署在自动驾驶和机器人控制等安全关键领域,还应用于推荐和行为分析等数据挖掘场景。虽然合规驱动的数据移除增强了隐私保护,但也开辟了一个此前未被识别的攻击面。我们提出了UBA-ORL(离线强化学习中的反学习激活后门攻击),这是首个针对离线强化学习的反学习激活后门攻击:在评估的设置中,该攻击在正常训练后受到显著抑制,而在合规驱动的删除(反学习)请求之后变得明显。UBA-ORL采用双样本机制:除了将触发器与恶意动作关联并在膨胀奖励下进行的后门轨迹(BD)之外,攻击者还注入共享相同触发器模式但保留良性动作且具有同样高奖励的伪装轨迹(CM)。在训练期间,BD和CM提供相互竞争的监督信号;当对CM子集发出合法删除请求后,残留的BD信号能够重新主导,按需重新激活后门。实证结果表明,在所评估的离线强化学习配置下,UBA-ORL实现了可控激活,而无触发器回报变化因配置而异,暴露了合规驱动的离线强化学习平台中一个此前被忽视的安全风险。我们敦促社区为合规的反学习服务开发联合的反学习前后审计机制。
英文摘要
Offline reinforcement learning (offline RL) enables policy learning from pre-collected static datasets without online exploration, and is increasingly deployed not only in safety-critical domains such as autonomous driving and robotic control but also in data-mining applications such as recommendation and behavior analysis. While compliance-driven data removal enhances privacy, it also opens a previously unrecognized attack surface. We introduce UBA-ORL (Unlearning-activated Backdoor Attack on Offline Reinforcement Learning), the first unlearning-activated backdoor attack for offline RL: in the evaluated settings, the attack is substantially suppressed after normal training and becomes pronounced after a compliance-driven deletion (unlearning) request. UBA-ORL employs a dual-sample mechanism: alongside backdoor trajectories (BD) that link a trigger to malicious actions under inflated rewards, the attacker injects camouflage trajectories (CM) sharing the same trigger pattern but preserving benign actions with equally high rewards. During training, BD and CM provide competing supervisory signals; upon a legitimate deletion request on the CM subset, the residual BD signal can re-dominate, reactivating the backdoor on demand. Empirical results show that UBA-ORL achieves controllable activation under the evaluated offline-RL configurations, while no-trigger return changes vary by configuration, exposing a previously overlooked security risk in compliance-driven offline RL platforms. We urge the community to develop joint pre-/post-unlearning auditing mechanisms for compliant unlearning services.
发表机构
- Huazhong University of Science and Technology(华中科技大学)
- Development and Education Center of the Cyberspace Administration of China(国家互联网信息办公室发展与教育中心)
机构由 AI 辅助整理,请以论文原文为准。