TASTE2:面向全双工语音交互的文本对齐语音建模与部署
TASTE2: Text-Aligned Speech Modeling and Deployment toward Full-Duplex Voice Interaction
浏览论文内容
中文总结 AI 辅助
TASTE2将TASTE扩展为增量对话栈,实现流式全双工语音交互,通过共享词表和增量解分词器提升性能,并首次系统表征副语言控制。
中文摘要 AI 辅助
全双工语音交互需要的不仅仅是话语级别的转换。它必须处理流式语音,管理话轮转换和打断,同时保留预训练的语言能力和声学副语言线索。我们探究TASTE(文本对齐语音分词与嵌入)是否为这一目标提供了可行路径。我们提出TASTE2,它将话语级别的TASTE转化为增量式对话栈。共享的文本-词元词汇表消除了词级平均,而模态对齐的对话训练在每个文本词元上预测一个连续的音频潜变量,无需交错异构词元流。增量式语音解分词器通过CosyVoice2实现流式合成。经过语音和对话训练后,TASTE2(合并)在LLaMA-Questions上达到56.3%,而Qwen2.5-7B指令文本-only参考为57.3%(准确率保留98.2%),TASTE2(直接)达到53.0%(保留92.4%)。我们构建了TASTE2 VoiceBot,它增量处理用户语音,流式传输合成音频,并在插入语音(barge-in)时停止生成。在Full-Duplex-Bench v1.0上,TASTE2和TASTE2 VoiceBot在保持高对话连贯性的同时很好地处理了打断。自然对话仍然具有挑战性,部署后经TensorRT加速,在两张NVIDIA RTX A6000上首次音频平均时间为2.701秒。最后,据我们所知,我们首次对基于TASTE的模型中的显式副语言控制进行了系统表征。在对话SFT后,快速语速作为跨策略的概念验证,而情感控制依赖于策略,其余属性仍然较弱。总之,这些结果确立了基于TASTE的建模作为实现全双工系统的实用途径,同时将自然对话鲁棒性、语音生成延迟和特征通用副语言控制确定为开放挑战。在线探索TASTE2。
英文摘要
Full-duplex voice interaction requires more than utterance-level conversion. It must process streaming speech, manage turn-taking and interruptions, while preserving pretrained linguistic competence and acoustic paralinguistic cues. We ask whether TASTE (Text-Aligned Speech Tokenization and Embedding) provides a viable path toward this goal. We present TASTE2, which transforms utterance-level TASTE into an incremental dialogue stack. A shared text-token vocabulary removes word-level averaging, while modality-aligned dialogue training predicts one continuous audio latent per text token without interleaving heterogeneous token streams. An incremental Speech Detokenizer enables streaming synthesis through CosyVoice2. After speech and dialogue training, TASTE2 (Merge) reaches 56.3% on LLaMA-Questions against a 57.3% Qwen2.5-7B Instruct text-only reference (98.2% accuracy retention), and TASTE2 (Direct) reaches 53.0% (92.4% retention). We build TASTE2 VoiceBot, which processes user speech incrementally, streams synthesized audio, and stops generation on barge-in. On Full-Duplex-Bench v1.0, TASTE2 and TASTE2 VoiceBot handle interruptions well while maintaining high conversational coherence. Natural conversation remains challenging, and deployed mean time to first audio is 2.701 s on two NVIDIA RTX A6000 after TensorRT acceleration. Finally, to our knowledge, we provide the first systematic characterization of explicit paralinguistic control in a TASTE based model. Fast speaking rate serves as a cross-strategy proof of concept after dialogue SFT, while emotion control is strategy dependent and the remaining attributes stay weak. Together, these results establish TASTE based modeling as a practical route toward full-duplex systems while identifying natural conversation robustness, speech generation latency, and feature general paralinguistic control as open challenges. Explore TASTE2 online.
发表机构
- MediaTek Research(联发科技研究院)
- National Taiwan University(台湾大学)
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。