arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VoiceChat-TTS:面向交互智能体的低延迟连续语音合成模型

VoiceChat-TTS: A Low-Latency Continuous Speech Synthesis Model for Interactive Agents

Edresson Casanova, Jaehyeon Kim, Mariana Graterol Fuenmayor, Shehzeen Hussain, Viacheslav Klimkov, Valentin Mendelev, Mikyas Desta, Paarth Neekhara, Piotr Zelasko, Chen Chen, Elena Rastorgueva, Ke Hu, Ankita Pasad, Xuesong Yang, Aya Alja'fari, Rajarshi Roy, Rohan Badlani, Jason Roche, Jason Li, Zhehuai Chen

arXiv 2608.13831首次发表:更新:

发表机构

NVIDIA Corporation(英伟达公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有语音模型延迟高、无法处理用户插话的问题,提出VoiceChat-TTS模型,由LLM文本令牌流驱动,支持语句中途中断且不重置KV缓存,在保证高语音质量的同时实现低延迟连续语音合成。

AI 中文摘要

口语对话是人机交互的自然形式,但大多数语音语言模型仍局限于轮次式运行,缺乏实时适应性,例如无法处理用户插话。近期的双工语音转语音、语音转文本模型通过替代多阶段流水线降低了延迟,但往往会牺牲语音质量,因为必须联合优化准确的自动语音识别(ASR)、中断处理和高保真合成。我们提出VoiceChat-TTS,一种面向交互智能体的低延迟、连续且可流式传输的文本到语音模型。VoiceChat-TTS由大语言模型(LLM)的文本令牌流直接驱动,通过控制令牌支持显式中断,且在无文本输入时生成静音。该模型在保持模块化和高语音质量的同时,实现了始终在线、响应迅速的语音生成,且支持语句中途中断,无需重置键值缓存(KV cache)。

英文摘要

Spoken dialogue is a natural form of human--computer interaction, yet most speech language models remain limited to turn-based operation and lack real-time adaptability, such as user barge-in. Recent duplex speech-to-speech and speech-to-text models reduce latency by replacing multi-stage pipelines, but often compromise speech quality because accurate ASR, interruption handling, and high-fidelity synthesis must be optimized jointly. We propose VoiceChat-TTS, a low-latency, continuous, and streamable text-to-speech model for interactive agents. VoiceChat-TTS is driven directly by LLM text-token streams, supports explicit interruption via control tokens, and produces silence when no textual input is available. The model enables always-on, responsive speech generation while preserving modularity and high speech quality, and it supports mid-utterance interruptions without resetting the KV cache.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑