噪声自适应流式音视频语音令牌增强用于鲁棒全双工口语对话模型
Noise Adaptive Streaming Audio-Visual Speech Token Enhancement for Robust Full-Duplex Spoken Dialogue Models
浏览论文内容
中文总结 AI 辅助
提出模块化流式音视频前端AV-STE,在语音LLM前从噪声音频和唇部视频恢复语义令牌,冻结下游模型,提升全双工对话鲁棒性,响应连贯性显著提高。
中文摘要 AI 辅助
全双工口语对话系统能够实现同时听与说,但其仅依赖音频的感知在背景噪声和重叠语音下常常失效,导致响应不连贯。近期的音视频对话方法表明,融入如唇部运动等视觉线索可提升在音频损坏情况下的鲁棒性。然而,现有方法往往需要调整大型口语对话模型本身以处理视觉输入,这要求昂贵的多模态训练。我们提出AV-STE,一种模块化流式音视频前端,在语音令牌到达语音大语言模型之前,从含噪音频和唇部视频中恢复受损的语义语音令牌。下游对话模型完全保持冻结,保留其预训练的对话能力。当与冻结的Moshi集成时,AV-STE在同数据集说话人干扰下将平均GPT-4o评判的响应连贯性从1.42提升至1.91,同时基本保持轮流发言行为。这些增益也迁移至域外的Seamless Interaction。
英文摘要
Full-duplex spoken dialogue systems enable simultaneous listening and speaking, but their audio-only perception often fails under background noise and overlapping speech, leading to incoherent responses. Recent audio-visual dialogue approaches show that incorporating visual cues such as lip movements improve robustness under audio corruption. However, existing approaches often adapt the large speech dialogue model itself to process visual input, requiring costly multimodal training. We propose AV-STE, a modular streaming audio-visual front-end that restores corrupted semantic speech tokens from noisy audio and lip video before they reach the speech LLM. The downstream dialogue model remains entirely frozen, preserving its pretrained conversational capabilities. When integrated with frozen Moshi, AV-STE improves average GPT-4o-judged response coherence from 1.42 to 1.91 under same-dataset speaker interference while largely preserving turn-taking behavior. Gains also transfer to out-of-domain Seamless Interaction.
发表机构
- KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。