发表机构
NVIDIA Corporation(英伟达公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有语音模型延迟高、无法处理用户插话的问题,提出VoiceChat-TTS模型,由LLM文本令牌流驱动,支持语句中途中断且不重置KV缓存,在保证高语音质量的同时实现低延迟连续语音合成。
AI 中文摘要
口语对话是人机交互的自然形式,但大多数语音语言模型仍局限于轮次式运行,缺乏实时适应性,例如无法处理用户插话。近期的双工语音转语音、语音转文本模型通过替代多阶段流水线降低了延迟,但往往会牺牲语音质量,因为必须联合优化准确的自动语音识别(ASR)、中断处理和高保真合成。我们提出VoiceChat-TTS,一种面向交互智能体的低延迟、连续且可流式传输的文本到语音模型。VoiceChat-TTS由大语言模型(LLM)的文本令牌流直接驱动,通过控制令牌支持显式中断,且在无文本输入时生成静音。该模型在保持模块化和高语音质量的同时,实现了始终在线、响应迅速的语音生成,且支持语句中途中断,无需重置键值缓存(KV cache)。
英文摘要
Spoken dialogue is a natural form of human--computer interaction, yet most speech language models remain limited to turn-based operation and lack real-time adaptability, such as user barge-in. Recent duplex speech-to-speech and speech-to-text models reduce latency by replacing multi-stage pipelines, but often compromise speech quality because accurate ASR, interruption handling, and high-fidelity synthesis must be optimized jointly. We propose VoiceChat-TTS, a low-latency, continuous, and streamable text-to-speech model for interactive agents. VoiceChat-TTS is driven directly by LLM text-token streams, supports explicit interruption via control tokens, and produces silence when no textual input is available. The model enables always-on, responsive speech generation while preserving modularity and high speech quality, and it supports mid-utterance interruptions without resetting the KV cache.