基于LLM的自动语音识别中的推理时目标说话人遗忘
Inference-Time Target Speaker Unlearning in LLM-Based Automatic Speech Recognition
- Indiana University(印第安纳大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出推理时目标说话人遗忘任务及轻量级注册条件门控模块,在冻结的语音LLM上动态实现新说话人退出转写,在AMI和AliMeeting上显著降低退出词准确率且不影响保留说话人性能。
AI中文摘要:
我们在一个完全端到端的框架中引入了目标说话人遗忘自动语音识别(TSU-ASR)任务,用于多说话人语音识别和说话人日志。给定一个多说话人话语和一组选择退出(opt-out)的说话人(他们不希望自己的语音被转写),该任务要求ASR系统转写除选择退出说话人之外的所有说话人,同时仍需指示这些说话人何时处于活跃状态。作为解决该任务的第一步,我们引入了一种新颖的轻量级注册条件门控(ECG)模块,可附加到冻结的双流语音大语言模型(LLM)上,使得在推理过程中能够动态地为新的选择退出说话人启用ASR,甚至包括那些在初始ECG训练阶段未见过的说话人。我们在AMI(英语)和AliMeeting(普通话)数据集上的实验表明,相应的选择退出词或字符的语音转写准确率分别从72.3%下降到48.2%,以及从73.6%下降到27.3%,而保留说话人的转写错误率基本保持不变。我们的方法为现代视频会议平台提供了一种实用解决方案,允许说话人动态地从自动化AI转写中选择退出,而无需强制离开会议会话,从而为每天可能数百万次的在线会议提供了一种保护隐私的接口。
英文摘要:
We introduce target-speaker unlearning ASR (TSU-ASR) task in a fully end-to-end framework for multi-speaker ASR and diarization. Given a multi-speaker utterance and a set of opt-out speakers who do not wish to have their speech transcribed, the task requires an ASR system to transcribe all speakers except the opt-out ones, while still indicating when those speakers are active. As a first step towards tackling this task, we introduce a novel, light-weight Enrollment-Conditioned Gating (ECG) module attachable to a frozen dual-stream speech LLM that enables ASR for new opt-out speakers dynamically during inference, even those who were not seen during initial ECG training phase. Our experiments on both AMI (English) and AliMeeting (Mandarin) datasets show that speech transcription accuracy for corresponding opt-out words or characters falls from 72.3% to 48.2% and from 73.6% to 27.3%, respectively, while retained speakers' transcription error rates maintain more or less the same. Our approach provides a practical solution for modern video conferencing platforms, allowing speakers to dynamically opt-out from automated AI transcriptions without forcefully leaving the meeting sessions, enabling a privacy-preserving interface for potentially millions of online meetings daily.