探索单一自回归LLM在同步和异步线索下统一目标语音提取
Exploring a Single Autoregressive LLM for Unified Target Speech Extraction across Synchronous and Asynchronous Cues
AI总结:
TSE-Omni利用单一自回归LLM,通过自注册的下一令牌预测统一处理同步和异步线索的目标语音提取,在干净和损坏视觉下均保持高性能,支持流式推理。
AI中文摘要:
目标语音提取(TSE)通常针对每种线索训练单独的提取器,而视觉线索系统往往需要与损坏匹配的训练才能在视觉帧损坏时保持鲁棒性。我们证明,一个自回归LLM骨干网络TSE-Omni可以同时服务于时间同步线索(唇部运动、伴随语音的手势)和异步线索(注册音频、文本)。TSE-Omni由下一令牌预测驱动:每一步从其自身过去的输出预测目标语音语义令牌,我们称之为自注册,形成由注册线索(异步音频或文本,或短视觉前缀)初始化的连续目标语音上下文。这实现了音视频补偿:模型在视觉完整时使用同步视觉,在视觉帧缺失时使用其令牌历史。在干净视觉下,TSE-Omni在VoxCeleb2和LRS3零样本上匹配强判别式和生成式基线(SpeechBERTScore分别为0.81和0.89),且DNSMOS更高。在相同的VoxCeleb2测试集上,经过2秒干净视觉起始后,移除剩余视觉帧使SpeechBERTScore保持在0.81。它在稀疏重叠和多说话人干扰下仍可用,并支持流式推理。项目页面:此https URL。
英文摘要:
Target speech extraction (TSE) typically trains a separate extractor per cue, and visual-cue systems often need corruption-matched training to remain robust under visual frame corruption. We show that one autoregressive LLM backbone, TSE-Omni, can serve both temporally synchronous cues (lip movements, co-speech gestures) and asynchronous cues (enrollment audio, text). TSE-Omni is driven by next-token prediction: each step predicts target speech semantic tokens from its own past outputs, which we term self-enrollment, forming a continuous target-speech context initialized by the enrollment cue (asynchronous audio or text, or a short visual prefix). This enables audio-visual compensation: the model uses synchronized visuals when intact and its token history when visual frames are missing. Under clean visuals, TSE-Omni matches strong discriminative and generative baselines (SpeechBERTScore 0.81 on VoxCeleb2 and 0.89 on LRS3 zero-shot) with higher DNSMOS. On the same VoxCeleb2 test set, after a 2 s clean visual start, removing the remaining visual frames leaves SpeechBERTScore at 0.81. It remains usable under sparse overlap and multi-speaker interference, and supports streaming inference. Project page: https://alexwxwu.github.io/tseomni-main/.