arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型进行同步语音翻译需要架构改变吗?一种前缀到前缀的数据驱动方法

Do LLMs Need Architectural Changes for Simultaneous Speech Translation? A Prefix-to-Prefix Data Driven Approach

Junkun Chen, Jian Xue, Ming Tang, Abdel Heba, Hoda Gholami, Ruchao Fan, Jinyu Li

arXiv 2607.13158首次发表:更新:

发表机构

Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究同步语音翻译中仅解码器语言模型面临的挑战,提出基于固定长度块、回退前缀及教师标记前缀到前缀目标的CSSEL-P2P方法,经实验其在可比延迟下提升了流质量,证明无需架构改变可有效实现同步语音翻译。

AI 中文摘要

同步语音翻译(SimulST)需要在严格的延迟约束下进行增量翻译,但对于仅解码器的语言模型系统来说仍然具有挑战性,因为上下文有限和跨语言重新排序。最近的方法通常引入架构改变或明确的读/写策略来控制输出时间,在分割边界不明确的对话语音中可能很脆弱。我们提出了一种简单的数据驱动替代方案:用于累积流解码的固定长度块,带有基于回退的提交前缀,以及带有有限等待的教师标记的前缀到前缀(P2P)目标进行微调,产生CSSEL-P2P,其中CSSEL是我们提出的分块流语音编码器语言模型。在我们的内部对话语音评估中,CSSEL-P2P在可比延迟(平均滞后0.15秒)下比CSSEL流基线的流质量提高了1.54 COMETKiwi,表明通过P2P监督无需架构改变即可实现有效的SimulST。

英文摘要

Simultaneous speech translation (SimulST) requires incremental translation under strict latency constraints, yet remains challenging for decoder-only LLM systems due to limited context and cross-lingual reordering. Recent approaches often introduce architectural changes or explicit read/write policies to control output timing, which can be brittle in conversational speech where segmentation boundaries are ambiguous. We present a simple data-driven alternative: fixed-length chunks for cumulative streaming decoding with a rewind-based committed prefix, and teacher-labeled prefix-to-prefix (P2P) targets with bounded waiting for fine-tuning, yielding CSSEL-P2P, where CSSEL is our proposed chunked streaming speech encoder LLM. In our in-house conversational speech evaluation, CSSEL-P2P improves streaming quality by +1.54 COMETKiwi over the CSSEL streaming baseline at comparable latency (+0.15s Average Lagging), suggesting effective SimulST without architectural changes via P2P supervision.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑