发表机构
Meta Platforms, Inc.(Meta平台公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对大型音频语言模型中的语音问答任务,提出多种机器遗忘策略(如梯度上升、任务算术和对齐微调),在降低高达80%隐私泄露率的同时,保持非私有任务性能接近不变。
AI 中文摘要
大型音频语言模型(LALMs)近期在语音理解和问答(QA)方面展现出强大的能力,但它们也从大规模训练数据中继承了隐私风险,包括对敏感信息的意外记忆。在本工作中,我们研究了LALMs中语音问答的机器遗忘,这一设置比先前基于文本的大型语言模型(LLMs)或自动语音识别(ASR)的工作更具挑战性,因为声学感知与事实知识之间存在紧密耦合。我们提出并评估了多种遗忘策略,包括梯度上升、任务算术以及基于对齐的微调方法(这些方法强制生成安全的拒绝响应),以在移除私有知识的同时保持核心能力的性能。通过在语音问答数据集上的大量实验,我们表明这些遗忘方法可以将隐私泄露率降低高达80%,同时在非私有语音问答和通用语音理解基准上保持接近中性的性能。
英文摘要
Large Audio-Language Models (LALMs) have recently shown strong capabilities in speech understanding and question answering (QA), but they also inherit privacy risks from large-scale training data, including the unintended memorization of sensitive information. In this work, we study machine unlearning for speech QA in LALMs, a setting that is more challenging than prior work on text-based Large Language Models (LLMs) or Automatic Speech Recognition (ASR) due to the tight coupling between acoustic perception and factual knowledge. We present and evaluate multiple unlearning strategies, including gradient ascent, task arithmetic, and alignment-based fine-tuning methods that enforce safe refusal responses, to remove private knowledge while still preserving performance on core capabilities. Through extensive experiments on speech QA datasets, we show that these unlearning methods can reduce the privacy leakage rate by up to 80% while maintaining near-neutral performance on non-private speech QA and general speech understanding benchmarks.