arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.09677cs.CL

X2-NativeCursor:面向增量文本流式编解码器TTS的原生词元文本进度跟踪

X2-NativeCursor: Native-Token Text Progress Tracking for Incremental-Text Streaming Codec TTS

Zehan Liu, Carl Chen, Rime Wen, Kaiqi Fu, Altman Lin, Shawn Qin, Lights Shi, Roy Gan, Hao Wang, Qian Wang

首次发表
浏览论文内容

中文总结 AI 辅助

提出X2-NativeCursor,一种轻量级观察器,利用原生语音词元在波形解码前跟踪增量TTS的文本进度,不改变生成器,显著降低对齐误差和实时因子。

中文摘要 AI 辅助

增量文本流式文本到语音合成(TTS)需要在线文本进度跟踪,以实现同步高亮、中断处理和对话历史更新。输入文本在语音播放之前就已到达,因此文本到达本身不能指示语音进度。现有的基于波形的对齐方法需要完整的音频,或在流式处理过程中增加声学处理。我们提出X2-NativeCursor,一种轻量级观察器,在波形解码之前从原生语音词元中跟踪进度,且不改变TTS生成器。其归一化计划将口语标签链接到原始文本片段。文本和原生词元编码器馈送给局部匹配器,该匹配器估计当前标签位置。一个单独的输出规则将可修正的位置估计转换为永不回退的光标。与自动参考相比,平均绝对误差为0.151个中文字符,前瞻80毫秒,而在线波形基线在320毫秒前瞻下的误差为1.253个字符。对齐实时因子也相对于该基线从0.3598降至0.0180。在第二个自动对齐参考下,较低的跟踪误差得以保持。我们在Qwen3-TTS上评估了X2-NativeCursor,并通过为每个骨干网络训练单独的观察器验证了其对CosyVoice2的适应性。代码可从此https URL公开获取。

英文摘要

Incremental-text streaming text-to-speech (TTS) needs online text progress tracking for synchronized highlighting, interruption handling, and dialogue-history updates. Input text arrives before it is spoken, so text arrival alone cannot indicate speech progress. Existing waveform-based alignment requires complete audio or adds acoustic processing during streaming. We propose X2-NativeCursor, a lightweight observer that tracks progress from native speech tokens before waveform decoding without changing the TTS generator. Its normalization plan links spoken labels to their original-text spans. Text and native-token encoders feed a local matcher that estimates the current label position. A separate output rule converts revisable position estimates into a cursor that never moves backward. Mean absolute error against an automatic reference is 0.151 Chinese characters with 80-ms lookahead, versus 1.253 characters with 320-ms lookahead for an online waveform baseline. Alignment real-time factor also decreases from 0.3598 to 0.0180 relative to this baseline. Lower tracking error is retained under a second automatic alignment reference. We evaluate X2-NativeCursor on Qwen3-TTS and validate its adaptation to CosyVoice2 by training a separate observer for each backbone. Code is publicly available at https://github.com/X-Square-Robot/X2Streaming-TTS.

发表机构

  • X Square Robot

机构由 AI 辅助整理,请以论文原文为准。

↑