arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越二元奖励:强化遗忘的奖励设计比较研究

Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning

Efstratios Zaradoukas, Davide Gabrielli, Bardh Prenkaj, Gjergji Kasneci

arXiv 2607.27968首次发表:更新:

发表机构

Technical University of Munich; Sapienza University of Rome; Munich Center for Machine Learning(慕尼黑工业大学; 罗马大学; 慕尼黑机器学习中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究在RUL框架下提出两种新型奖励函数,经RWKU基准实验,其遗忘效率较二元奖励提升,速度快达3倍且保留模型效用,表明奖励设计是机器遗忘效率的关键驱动因素。

AI 中文摘要

机器遗忘旨在从训练好的语言模型中有选择地移除特定知识而无需完全重新训练,在GDPR和欧盟AI法案等隐私法规下这一需求日益增长。近期研究将遗忘重新表述为可验证奖励强化学习(RLVR)问题,其中模型针对直接从其输出计算的可验证奖励进行优化。然而,现有方法依赖稀疏二元奖励,仅提供极少学习信号,仅表明是否避免了禁止内容,限制了收敛速度。本文在强化遗忘(RUL)框架内研究奖励设计如何影响遗忘效率,引入原则性奖励分解框架,将可验证性与稀疏性解耦,并提出两种新奖励函数:指数奖励,基于禁止概念出现次数提供分级惩罚;受PageRank启发的奖励,按语义重要性加权惩罚。我们在现实世界知识遗忘(RWKU)基准上进行实验,证明两种奖励均持续优于二元设置,同时达到相似遗忘性能的速度快达3倍,并保留了通用模型效用。我们的结果表明,奖励设计是遗忘效率的关键驱动因素,为可扩展且高效的机器遗忘提供了实用路径。

英文摘要

Machine unlearning seeks to selectively remove specific knowledge from trained language models without full retraining, a growing necessity under privacy regulations such as GDPR and the EU AI Act. Recent work has reformulated unlearning as a Reinforcement Learning with Verifiable Rewards (RLVR) problem, where models are optimized against verifiable rewards computed directly from their outputs. However, existing methods rely on sparse binary rewards that provide minimal learning signal, indicating only whether forbidden content was avoided, and limiting convergence speed. In this paper, we study how reward design affects unlearning efficiency within the Reinforcement Unlearning (RUL) framework. We introduce a principled reward decomposition framework that decouples verifiability from sparsity, and propose two new reward functions: an exponential reward that provides graded penalties based on the count of forbidden-concept occurrences, and a PageRank inspired reward that weights penalties by semantic importance. We conduct experiments on the Real World Knowledge Unlearning (RWKU) benchmark, demonstrating that both rewards consistently outperform the binary setting, while reaching similar forgetting performance up to $3\times$ faster and preserving general model utility. Our results show that reward design is a key driver of unlearning efficiency offering a practical path toward scalable and efficient machine unlearning.

CommentsAccepted to WIPE-OUT 2 @ ECML-PKDD 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑