arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10878cs.CLeess.AS

X2-Turn:用于联合流式ASR与轮次状态预测的帧同步双头建模

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction

Kaiqi Fu, Rime Wen, Altman Lin, Shawn Qin, Roy Gan, Hao Wang, Qian Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出X2-Turn方法,基于预训练Voxtral Realtime模型构建帧同步双头结构,联合预测ASR token与细粒度轮次状态,在中英双语Easy-Turn测试集上验证了其轮次转换检测的准确性与低延迟优势。

中文摘要 AI 辅助

准确且响应迅速的轮次转换对于口语对话系统至关重要,这类系统必须实时区分用户打断、应忽略的反馈信道以及话语的完成情况。现有模块化方法通常在话语或固定块级别优化轮次状态预测,与连续轮次状态估计存在不匹配,且往往依赖辅助ASR模型,这限制了响应速度并增加了整体系统复杂度。因此,我们提出X2-Turn,一种基于延迟流式建模的帧同步轮次状态预测方法。具体而言,我们在预训练的Voxtral Realtime模型基础上,引入了一个帧同步轮次状态头,该头与ASR头在共享的流式表示上并行运行,共同在帧级别预测ASR token和细粒度轮次状态。我们在中英双语Easy-Turn测试集上评估了该方法,结果表明其在实现准确轮次转换检测的同时保持低延迟方面具有有效性。

英文摘要

Accurate and responsive turn-taking is essential for spoken dialogue systems, which must distinguish in real time between user interruptions, backchannels that should be ignored, and the completion of an utterance. Prior modular approaches typically optimize turn state prediction at the utterance or fixed-chunk level, creating a mismatch with the continuous turn state estimate, and often depend on an auxiliary ASR model, which limits responsiveness and increases overall system complexity. Therefore, we present X2-Turn, a frame-synchronous turn state prediction method via delayed-stream modeling. Specifically, building on the pretrained Voxtral Realtime model, we introduce a frame-synchronous turn state head that operates in parallel with the ASR head on shared streaming representations, jointly predicting ASR tokens and fine-grained turn states at the frame level. Experiments on bilingual EasyTurn and Full-Duplex-Bench demonstrate that the proposed method achieves an effective trade-off between turn state accuracy and decision latency.

↑