arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2609.24267eess.AS

StreamTN:面向对话系统中流式TTS的低延迟中文文本规范化模型

StreamTN: A Low-Latency Streaming Chinese Text Normalization Model for Streaming TTS in Dialogue Systems

  • Northwestern Polytechnical University(西北工业大学)
  • Shenzhen Pimei Technology Co., Ltd.(深圳皮米科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Wenhao Li, Jinrui Liang, Haoyu Zhang, Jingbin Hu, Xiaming Ren, Hanke Xie, Huakang Chen, Chengyou Wang, Dake Guo, Linhan Ma, Su Feng, Houdun Liu, Yunxiang Chen, Lei Xie

AI总结:

StreamTN是一种基于Qwen3-0.6B的双轨流式中文文本规范化模型,通过任务微调实现低延迟、高准确率,并引入新基准评估,适用于对话系统TTS。

AI中文摘要:

文本到语音(TTS)是以大型语言模型(LLM)为中心的语音对话系统(SDS)中提供语音响应的关键模块。为确保TTS合成的准确性,LLM生成的响应必须通过文本规范化(TN)模块转换为TTS可读格式,这在实时SDS场景中对低延迟提出了严格要求。现有的TN解决方案主要基于规则,依赖人工工程,且对未见模式泛化能力差。尽管LLM本身可以通过提示工程执行TN,但它面临关键限制:由于非流式处理导致的高首令牌延迟、幻觉风险,以及当核心LLM模块仅针对TN进行微调时智能或推理能力下降。为解决这些挑战,我们提出了StreamTN,一种基于LLM的轻量级中文流式TN模型。StreamTN基于Qwen3-0.6B构建,采用双轨流式框架,其中输入令牌和输出令牌在两条并行轨道上处理,无需复杂提示即可实现低延迟实时推理。此外,任务特定的微调相比基于规则的系统和大规模通用LLM,带来了更优的TN性能和更少的幻觉。我们还引入了一个涵盖多种文本场景的TN基准,为语音对话系统中的语音生成提供了全面的评估标准。实验证明了StreamTN在准确性和推理延迟方面的有效性。

英文摘要:

Text-to-Speech (TTS) is an essential module that provides spoken responses in a spoken dialogue system (SDS) centered on a large language model (LLM). To ensure accurate TTS synthesis, responses generated by an LLM must be converted into TTS-readable formats via a Text Normalization (TN) module, imposing strict low-latency requirements in real-time SDS scenarios. Existing TN solutions are largely rule-based, rely on manual engineering, and generalize poorly to unseen patterns. Although an LLM itself can perform TN through prompt engineering, it faces key limitations: high first-token latency due to non-streaming processing, hallucination risks, and degraded intelligence or reasoning when the core LLM module is fine-tuned solely for TN. To address these challenges, we propose StreamTN, a lightweight LLM-based Chinese streaming TN model. Built on Qwen3-0.6B, StreamTN employs a dual-track streaming framework in which input tokens and output tokens are processed on two parallel tracks, enabling low-latency real-time inference without complex prompting. Moreover, task-specific fine-tuning yields superior TN performance and fewer hallucinations than rule-based systems and general-purpose LLMs. We also introduce a TN benchmark that spans diverse text scenarios, providing a comprehensive evaluation standard for speech generation in spoken dialogue systems. Experiments demonstrate the effectiveness of StreamTN in accuracy and inference latency.

补充信息

↑