StreamTN:面向对话系统中流式TTS的低延迟中文文本规范化模型
StreamTN: A Low-Latency Streaming Chinese Text Normalization Model for Streaming TTS in Dialogue Systems
- Northwestern Polytechnical University(西北工业大学)
- Shenzhen Pimei Technology Co., Ltd.(深圳皮米科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
StreamTN是一种基于Qwen3-0.6B的双轨流式中文文本规范化模型,通过任务微调实现低延迟、高准确率,并引入新基准评估,适用于对话系统TTS。
AI中文摘要:
文本到语音(TTS)是以大型语言模型(LLM)为中心的语音对话系统(SDS)中提供语音响应的关键模块。为确保TTS合成的准确性,LLM生成的响应必须通过文本规范化(TN)模块转换为TTS可读格式,这在实时SDS场景中对低延迟提出了严格要求。现有的TN解决方案主要基于规则,依赖人工工程,且对未见模式泛化能力差。尽管LLM本身可以通过提示工程执行TN,但它面临关键限制:由于非流式处理导致的高首令牌延迟、幻觉风险,以及当核心LLM模块仅针对TN进行微调时智能或推理能力下降。为解决这些挑战,我们提出了StreamTN,一种基于LLM的轻量级中文流式TN模型。StreamTN基于Qwen3-0.6B构建,采用双轨流式框架,其中输入令牌和输出令牌在两条并行轨道上处理,无需复杂提示即可实现低延迟实时推理。此外,任务特定的微调相比基于规则的系统和大规模通用LLM,带来了更优的TN性能和更少的幻觉。我们还引入了一个涵盖多种文本场景的TN基准,为语音对话系统中的语音生成提供了全面的评估标准。实验证明了StreamTN在准确性和推理延迟方面的有效性。
英文摘要:
Text-to-Speech (TTS) is an essential module that provides spoken responses in a spoken dialogue system (SDS) centered on a large language model (LLM). To ensure accurate TTS synthesis, responses generated by an LLM must be converted into TTS-readable formats via a Text Normalization (TN) module, imposing strict low-latency requirements in real-time SDS scenarios. Existing TN solutions are largely rule-based, rely on manual engineering, and generalize poorly to unseen patterns. Although an LLM itself can perform TN through prompt engineering, it faces key limitations: high first-token latency due to non-streaming processing, hallucination risks, and degraded intelligence or reasoning when the core LLM module is fine-tuned solely for TN. To address these challenges, we propose StreamTN, a lightweight LLM-based Chinese streaming TN model. Built on Qwen3-0.6B, StreamTN employs a dual-track streaming framework in which input tokens and output tokens are processed on two parallel tracks, enabling low-latency real-time inference without complex prompting. Moreover, task-specific fine-tuning yields superior TN performance and fewer hallucinations than rule-based systems and general-purpose LLMs. We also introduce a TN benchmark that spans diverse text scenarios, providing a comprehensive evaluation standard for speech generation in spoken dialogue systems. Experiments demonstrate the effectiveness of StreamTN in accuracy and inference latency.