对遗忘算法进行压力测试
Stress Testing Unlearning Algorithms
AI总结:
该研究针对现有大型语言模型遗忘基准的缺陷,提出扩展WMDP的WMDP++基准,以针对性提取遗忘信息和评估边界问题性能,实现对遗忘算法更严格的压力测试。
AI中文摘要:
近期,机器遗忘(即从模型中移除特定训练数据的影响)受到越来越多的关注。在大型语言模型(LLMs)中,由于输入和输出的模糊性,遗忘过程尤其具有挑战性。因此,严格的评估对于评估安全性和实用性、推动遗忘方法的进步至关重要。我们发现现有遗忘基准存在两个关键缺陷:(1)它们未主动测试遗忘的信息是否仍可被强行提取;(2)它们无法评估在边界问题(即与遗忘内容语义接近的良性查询)上的性能保持情况。在此,我们引入WMDP++,它是WMDP的扩展版本,通过整合对遗忘信息的针对性提取以及对边界问题的系统性评估来解决这些缺陷。WMDP++为评估大型语言模型的遗忘提供了更严格且更具信息性的基准。
英文摘要:
Recently, machine unlearning, the removal of specific training data influence from a model, has gained increasing attention. In large language models (LLMs), unlearning is particularly challenging due to the ambiguity of inputs and outputs. Con- sequently, rigorous evaluation is critical for assessing both safety and utility, and for driving progress in unlearning meth- ods. We identify two key shortcomings in existing unlearning benchmarks: (1) they do not actively test whether unlearned information can still be forcibly extracted, and (2) they fail to evaluate performance preservation on boundary questions, be- nign queries that are semantically close to the unlearned con- tent. Here we introduce WMDP++, an extension of WMDP that addresses these gaps by incorporating targeted extrac- tion of unlearned information and systematic evaluation on boundary questions. WMDP++ provides a more stringent and informative benchmark for evaluating unlearning in LLMs.