发表机构
Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究同步语音翻译中仅解码器语言模型面临的挑战,提出基于固定长度块、回退前缀及教师标记前缀到前缀目标的CSSEL-P2P方法,经实验其在可比延迟下提升了流质量,证明无需架构改变可有效实现同步语音翻译。
AI 中文摘要
同步语音翻译(SimulST)需要在严格的延迟约束下进行增量翻译,但对于仅解码器的语言模型系统来说仍然具有挑战性,因为上下文有限和跨语言重新排序。最近的方法通常引入架构改变或明确的读/写策略来控制输出时间,在分割边界不明确的对话语音中可能很脆弱。我们提出了一种简单的数据驱动替代方案:用于累积流解码的固定长度块,带有基于回退的提交前缀,以及带有有限等待的教师标记的前缀到前缀(P2P)目标进行微调,产生CSSEL-P2P,其中CSSEL是我们提出的分块流语音编码器语言模型。在我们的内部对话语音评估中,CSSEL-P2P在可比延迟(平均滞后0.15秒)下比CSSEL流基线的流质量提高了1.54 COMETKiwi,表明通过P2P监督无需架构改变即可实现有效的SimulST。
英文摘要
Simultaneous speech translation (SimulST) requires incremental translation under strict latency constraints, yet remains challenging for decoder-only LLM systems due to limited context and cross-lingual reordering. Recent approaches often introduce architectural changes or explicit read/write policies to control output timing, which can be brittle in conversational speech where segmentation boundaries are ambiguous. We present a simple data-driven alternative: fixed-length chunks for cumulative streaming decoding with a rewind-based committed prefix, and teacher-labeled prefix-to-prefix (P2P) targets with bounded waiting for fine-tuning, yielding CSSEL-P2P, where CSSEL is our proposed chunked streaming speech encoder LLM. In our in-house conversational speech evaluation, CSSEL-P2P improves streaming quality by +1.54 COMETKiwi over the CSSEL streaming baseline at comparable latency (+0.15s Average Lagging), suggesting effective SimulST without architectural changes via P2P supervision.