arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19442cs.LGcs.AI

作为分布恢复的遗忘学习:一项可控的反事实研究、一个经过验证的选择性筛选以及无预言机认证的局限性

Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification

Sen Yang, Yuen-Hei Yeung

AI总结:

研究机器遗忘学习评估标准,发现现有标准有局限。重新定义遗忘学习为恢复到匹配参考,审核相关标准,介绍多种测试及结果,包括未使用筛选、损伤相对重新校准等,指出仅前向认证不可靠,提出对实际方法的实证选择性测试及可识别性定理。

AI中文摘要:

机器遗忘学习通常通过在训练探针上匹配重新训练的预言机来评估。在一个具有匹配重新训练参考的可控一次性测试平台中,我们发现该标准可能有利于保留未使用知识的方法:它评定为合格的候选者遗忘事实的得分比从未学习的水平低2.82奈特(聚类置信区间[-3.16, -2.48])。我们将遗忘学习重新定义为恢复到匹配参考,并在跨越五个开放架构家族的45个模型种子单元中审核无预言机筛选和证书式标准。参考本身证伪了绝对保留/往返证书:注入模型按构造保留了保留集,但在45个单元中的41个未通过固定保留阈值,在45个单元中的31个未通过自身往返测试,而参考仅在45个单元中的1个完全通过认证。基于基础的未使用筛选作为选择性必要测试仍然很强:在一个密封挑战套件上,它在45个单元中的45个拒绝了注入模型,在45个单元中的44个接受了参考,并部分检测到实体路由抑制(在其中45个单元中的35个);它是一个具有测量灵敏度的必要测试,而非充分性证书。基于参考自身操作点的损伤相对重新校准在45个单元中的15个认证了一个小子集;在不放弃的情况下,其选择在其优化轴上位于重新训练噪声(0.80奈特)范围内,而常见的训练探针标准则相差5.17奈特(这是一个支持性比较,而非直接基准测试)。固定幅度的逻辑抑制攻击在45个单元中的12个击败了完整的前向测试组,因此仅前向认证是不可靠的;我们的方法是对实际产生的方法进行的实证选择性测试。一个可识别性定理界定了哪些事实根本允许无预言机遗忘阈值,其中TOFU是预测的边界情况。

英文摘要:

Machine unlearning is commonly evaluated by matching a retrained oracle on trained probes. In a controlled nonce-fact testbed with a matched retraining reference, we find this criterion can favor methods that retain held-out knowledge: candidates it rates adequate score held-out forget facts $-2.82$ nats below the never-learned level (cluster CI $[-3.16,-2.48]$). We recast unlearning as restoration to the matched reference and audit oracle-free screens and certificate-style criteria across 45 model-seed cells spanning five open architecture families. The reference itself falsifies an absolute retain/round-trip certificate: the injected model, which retains the retain set by construction, fails the fixed retain threshold in 41/45 cells and its own round trip in 31/45, and the reference fully certifies in only 1/45. A base-anchored held-out screen remains strong as a selective necessary test: on a sealed challenge suite it rejects the injected model in 45/45 cells, accepts the reference in 44/45, and partially detects entity-routing suppression (35/45); it is a necessary test with measured sensitivity, not a sufficiency certificate. A damage-relative recalibration anchored to the reference's own operating point certifies a small subset in 15/45 cells; where it does not abstain, its picks lie within retraining noise (0.80 nats) on the axes it optimizes, while the common trained-probe criterion sits 5.17 nats away (a supporting comparison, not a head-to-head benchmark). A fixed-magnitude logit-suppression attack defeats the full forward battery in 12/45 cells, so forward-only certification is not sound; our method is an empirical selective test for methods-as-produced. An identifiability theorem delimits which facts admit an oracle-free forget threshold at all, with TOFU as the predicted boundary case.

↑