发表机构
INESC-ID; Instituto Superior Técnico, Universidade de Lisboa; LTI, Carnegie Mellon University(INESC-ID; 里斯本大学高等技术学院; 卡内基梅隆大学语言技术研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文首次系统评估机器遗忘在自动语音识别中的适用性,发现梯度上升算法效果最佳,而复杂方法易过度遗忘,且顺序/同时遗忘性能劣于单主体遗忘,需改进评估方法。
AI 中文摘要
机器遗忘(MU)为遵守“被遗忘权”法规提供了一条途径。尽管MU在语音任务中受到越来越多的关注,但在自动语音识别(ASR)领域仍 largely 未被探索。在本工作中,我们调查现有MU算法和评估工具是否适用于ASR。我们将多种MU技术应用于一个ASR模型,评估单主体遗忘的隐私-效用权衡,然后在顺序遗忘和同时遗忘下评估最佳算法。结果表明,基于梯度上升的算法实现了强大的效用-隐私权衡,而更复杂的方法过度遗忘样本,使其更容易被识别为已遗忘。这表明基于简单成员推断攻击的标准隐私评估不足以可靠地评估遗忘成功,促使改进ASR中MU的评估方法。最后,我们表明顺序遗忘和同时遗忘在隐私和效用方面均不如单主体遗忘,强调了需要更适合这些场景的遗忘构建。
英文摘要
Machine unlearning (MU) offers a path to compliance with "right to be forgotten" regulations. While MU has received increasing attention for speech tasks, it remains largely unexplored for Automatic Speech Recognition (ASR). In this work, we investigate whether existing MU algorithms and evaluation tools are suitable for ASR. We apply several MU techniques to an ASR model, evaluating privacy-utility trade-offs for single-subject unlearning, then assess the best algorithm under sequential and simultaneous unlearning. Results show that gradient ascent-based algorithms achieve strong utility-privacy trade-offs, whereas more complex approaches over-unlearn samples, making them easier to identify as unlearned. This suggests standard privacy evaluations based on simple Membership Inference attacks are insufficient to reliably assess unlearning success, motivating improved evaluation methods for MU in ASR. Finally, we show that both sequential and simultaneous unlearning yield worse privacy and utility than single-subject unlearning, underscoring the need for unlearning constructions better suited to these settings.
CommentsSubmitted to ICASSP 2027