SimulS2ST-Omni:通过显式轨迹监督实现数据高效的流式语音到语音翻译
SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision
查看机构详情
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究长格式流式语音到语音翻译,提出仅用约2000小时配对数据的训练方法,核心是联合文本-代码轨迹监督,双流分解减轻干扰,在质量-延迟权衡上表现出色,与先进闭源系统相当。
中文摘要 AI 辅助
长格式流式语音到语音翻译(S2ST)需要在严格的延迟约束下进行增量、无界翻译。现有方法通常受限于句子级监督或需要大量配对的S2ST监督。我们引入一种训练方法,仅使用约2000小时的配对跨语言S2ST数据,在辅助监督之上,使语音语言模型用于句子级和长格式流式S2ST。以辅助多任务训练为基础,即使配对S2ST预算减少90%,我们的方法仍保持稳健。核心贡献是联合文本-代码轨迹监督,消除对单独不稳定语音侧发射控制器的需求。双流Thinker--Talker分解通过解耦语言推理与密集声学预测减轻模态干扰,显著优于统一解码器基线。最后,我们的系统在RealSI和ACL60/60-dev上实现了极具竞争力的质量-延迟权衡,在ASR-BLEU上与LiveInterpret~2.0等最先进的闭源S2ST系统相当。
英文摘要
Long-form streaming speech-to-speech translation (S2ST) requires incremental, unbounded translation under strict latency constraints. Existing methods typically suffer from sentence-bounded supervision or demand massive paired-S2ST supervision. We introduce a training recipe enabling a speech language model for sentence-level and long-form streaming S2ST using only $\sim$2k hours of paired cross-lingual S2ST data, layered atop auxiliary supervision. Anchored by auxiliary multitask training, our approach remains robust even when the paired-S2ST budget itself is reduced by 90\%. Our core contribution, joint text-code trajectory supervision, schedules target text and acoustic semantic codes as a unified commitment path, eliminating the need for separate, unstable speech-side emission controllers. Furthermore, our two-stream Thinker--Talker factorization significantly outperforms unified-decoder baselines by decoupling linguistic reasoning from dense acoustic prediction to mitigate modality interference. Finally, our system achieves highly competitive quality-latency trade-offs on RealSI and ACL60/60-dev, matching state-of-the-art, closed-source S2ST systems such as LiveInterpret~2.0 on ASR-BLEU.