在串联语音到语音模型中通过随机引导学习自然对话行为
Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance
查看机构详情
- Sakana AI
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对串联语音到语音模型训练中引导数据缺失问题,提出随机中间引导方法,直接从语料库采样引导,提升响应质量与自然对话行为,优于基线模型。
中文摘要 AI 辅助
串联语音到语音架构将响应式语音前端与异步文本后端相结合。在KAME中,大型语言模型(LLM)作为后端,在用户仍在说话时向前端提供候选响应作为引导。普通对话录音捕获最终响应,但不捕获后端在用户话语期间会提供的引导。使用模拟器LLM生成缺失的引导,在基于真实对话训练时会增加大量的数据准备开销。我们提出随机中间引导,直接从对话语料库中推导引导,而非模拟后端LLM行为。在训练期间,目标响应提供信息丰富的引导,而随机采样的响应则在话语期间提供可能不相关的更新。这种组合旨在教导前端有选择性地使用后端信息。在合成对话上,使用此方法训练的KAME达到了与LLM生成和基于相似性的基线相当的响应质量。在3.8k小时的真实对话上训练,相比合成数据KAME,改善了平滑的轮流发言和音频评判的自然度,同时保持了对Moshi的响应质量优势。这些结果表明,随机引导提供了一条实用途径,将串联模型的响应质量优势与从真实语音中学习的自然对话行为相结合。
英文摘要
Tandem speech-to-speech architectures couple a responsive speech frontend with an asynchronous text backend. In KAME, a large language model (LLM) serves as the backend, supplying candidate responses as guidance to the speech frontend while the user is still speaking. Ordinary conversation recordings capture the eventual response but not the guidance the backend would supply during the user's utterance. Generating the missing guidance with a simulator LLM adds substantial data-preparation overhead when training on real conversations. We propose randomized intermediate guidance, which derives guidance directly from the conversation corpus rather than simulating backend LLM behavior. During training, target responses provide informative guidance, while randomly sampled responses provide potentially irrelevant updates during the utterance. This combination aims to teach the frontend to use backend information selectively. On synthetic dialogues, KAME trained with this recipe achieves response quality comparable to that of the LLM-generated and similarity-based baselines. Training on 3.8k hours of real conversations improves smooth turn-taking and audio-judge naturalness over synthetic-data KAME while retaining a response-quality advantage over Moshi. These results show that randomized guidance offers a practical route to combining the response-quality benefits of tandem models with natural conversational behavior learned from real speech.