arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于GRPO的大语言模型遗忘中奖励规范与基准可靠性的实证研究

An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning

Rubén Balbastre, Juan Manuel Orduña, Mariano Pérez

arXiv 2608.17804首次发表:更新:

发表机构

University of Valencia(瓦伦西亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对基于GRPO的LLM遗忘,探究奖励规范与基准可靠性问题,对比四种奖励设计开展实验,发现优化成功与行为遗忘并不等价,并分析了分歧的来源。

AI 中文摘要

实用的大语言模型(LLM)遗忘通常通过两个目标评估:抑制目标特定知识,以及保留非目标效用。在生成式问答场景中,这留下了第三种未明确的行为:当与目标相关的提示允许在不发生目标特定泄漏的情况下给出更宽泛的答案时,模型应按该级别回答,而非泄漏、规避或弃权(不执行)。我们在受控的LoRA-GRPO RWKU设置中研究这一规范问题,对比四种奖励设计,涵盖词汇抑制、反弃权塑造、基于 rubric 的宽泛回答以及明确的弃权对比,设置含与不含 SFT 预热两种情况。实验表明,优化成功不等同于行为遗忘:RWKU 遗忘分数、保留的完成审计、终端训练-部署审计及训练动态可指向不同结论。我们将这些分歧归因于奖励黑客端点、GRPO 的策略支持限制、未捕捉端点变化的基准探测,以及优化过程中可选择语义泄漏低的宽泛主题回答的奖励。

英文摘要

Practical LLM unlearning is usually evaluated through two objectives: suppress target-specific knowledge and preserve non-target utility. In generative QA, this leaves a third behavior underspecified: when a target-adjacent prompt admits a broader answer without target-specific leakage, the model should answer at that level rather than leak, evade, or refuse. We study this specification problem in a controlled LoRA-GRPO RWKU setting, comparing four reward designs that span lexical suppression, anti-refusal shaping, rubric-based broad answering, and an explicit refusal contrast, with and without SFT warm-up. The experiments show that optimization success is not equivalent to behavioral unlearning: RWKU forget scores, held-out completion audits, and training dynamics can point to different conclusions. We trace these disagreements to reward-hacking endpoints, policy-support limits in GRPO, benchmark probes that miss endpoint changes, and a rubric reward that selects broad-topic answering with low semantic leakage under held-out evaluation.

Comments29 pages, 5 figures. Code and artifacts linked in the paper. v2: Extended the held-out evaluation to include broad-topic helpfulness, replacing the terminal-training rollout analysis

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑