arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

全双工语音大语言模型中虚假起始的因果分析与缓解

Causal Analysis and Mitigation of Spurious Onsets in Full-Duplex Speech LLMs

Kento Nishi

arXiv 2609.13445首次发表:更新:

发表机构

Massachusetts Institute of Technology; Comcast Applied AI Research, Speech AI Team(麻省理工学院; 康卡斯特应用人工智能研究,语音人工智能团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对全双工语音LLM在用户沉默时虚假起始的问题,提出基于因果反事实的推理时方法,抑制虚假起始同时保留真实响应,无需重训练且实时运行。

AI 中文摘要

语音到语音的大语言模型,如Moshi及其衍生模型PersonaPlex,可以通过全双工生成同时进行听和说。然而,在用户长时间沉默期间,它们可能不恰当地开始说话:在数字零输入下,Moshi和PersonaPlex分别在12/40和11/40个五分钟的延续中发起语音。是什么导致了这种虚假语音?我们研究了两个假设:要么重复采样尽管起始概率持续较低仍选择语音,要么以模型的非语音输出为条件导致起始概率突然飙升。我们发现,在每次观察到的起始处,语音概率在80毫秒的帧内飙升超过九个数量级,支持后一个假设。然后,为了在不阻止真实响应的前提下抑制这些起始,我们提出了一个因果反事实问题:模型是在响应用户语音,还是如果前面的用户输入被静音,其下一个词元分布仍会保持相似?据此,我们抑制在该干预下分布变化很小的起始。在每模型40个保留试验中,使用真实麦克风噪声,我们的方法抑制了Moshi的13/13和PersonaPlex的9/9个虚假起始,同时保留了每模型40/40个真实响应。我们的推理时方法无需重新训练,实时运行,95百分位决策时间低于61毫秒,在80毫秒帧预算内。我们的代码可在以下网址获取:https://this https URL。

英文摘要

Speech-to-speech LLMs like Moshi, and its derivative PersonaPlex, can listen and speak concurrently through full-duplex generation. However, they can begin speaking inappropriately during prolonged user silence: under digital-zero input, Moshi and PersonaPlex initiate speech in 30% and 27.5% of five-minute continuations, respectively. What causes this spurious speech? We investigate two hypotheses: either repeated sampling selects speech despite persistently low onset probabilities, or self-conditioning on nonspeech outputs causes an abrupt spike in onset probability. We find that, at every observed onset, speech probability spikes by over nine orders of magnitude in one 80-ms frame, supporting the latter hypothesis. Then, to suppress these onsets without blocking genuine responses, we ask a causal counterfactual question: is the model responding to user speech, or would its next-token distribution remain similar if the preceding user input were muted? Accordingly, we suppress onsets whose distributions change little under this intervention. Under realistic microphone noise, our method suppresses spurious onsets, while preserving genuine responses: one-sided 95% lower confidence bounds are 98.68% and 98.82% for Moshi, and 96.90% and 99.25% for PersonaPlex. Our inference-time method runs in real-time without retraining, with 95th-percentile decision time below 61 ms, within the 80-ms frame budget. Our code is available at https://github.com/KentoNishi/icassp27-spurious-onsets.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑