arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

X2Streaming-ASR:不确定时等待,就绪时输出——面向流式语音识别

X2Streaming-ASR: wait when uncertain, emit when ready for streaming ASR

Zhiwei Lin, Kaiqi Fu, Rime Wen, Zehan Liu, Shawn Qin, Roy Gan, Hao Wang

arXiv 2609.08672首次发表:更新:

发表机构

X Square Robot(X Square机器人公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

X2Streaming-ASR通过三阶段训练将流式识别分解为提交时机与内容,在AISHELL和WenetSpeech上以27-84毫秒延迟实现最佳流式CER。

AI 中文摘要

面向实时语音智能体和全双工对话的流式自动语音识别(ASR)必须提供低提交延迟的准确部分转录。现有系统通常使用固定块大小、前瞻或目标延迟,或鼓励在估计的声学边界附近进行输出。这些方法并未在单遍、硬提交约束下直接优化每个输出位置应使用多少额外上下文。我们提出X2Streaming-ASR,将流式识别分解为何时提交和提交什么。其三阶段训练流程首先建立流式识别能力,然后使用自动探测的轨迹对提交策略进行热启动,最后使用基于字符级、分段分配的组相对奖励来优化策略,以兼顾识别准确性和延迟。在AISHELL-1/2/3和WenetSpeech上,相对于强制对齐的字符端点,X2Streaming-ASR实现了27-84毫秒的平均字符级提交延迟,而评估的流式基线为409-585毫秒。在AISHELL-1和AISHELL-3上,它以显著更低的延迟取得了评估系统中最佳的流式CER。

英文摘要

Streaming automatic speech recognition (ASR) for real-time voice agents and full-duplex dialogue must provide accurate partial transcripts with low commit latency. Existing systems commonly use a fixed chunk size, look-ahead, or target delay, or encourage emissions near estimated acoustic boundaries. These approaches do not directly optimize how much additional context to use at each output position under a single-pass, hard-commit constraint. We propose X2Streaming-ASR, which decomposes streaming recognition into when to commit and what to commit. Its three-stage training procedure first establishes streaming recognition ability, then warm-starts the commit policy with automatically probed trajectories, and finally refines the policy using character-level, segment-assigned group-relative rewards for recognition accuracy and latency. Across ten Chinese and English test sets, X2Streaming-ASR attains the lowest mean commit latency relative to forced-aligned endpoints, 32--109ms on Chinese characters and 12--85ms on English words, while recognition accuracy remains comparable to existing systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑