学习在部分语音中何时提交,用于端到端同声传译
Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation
浏览论文内容
中文总结 AI 辅助
本研究通过前缀监督适配语音语言模型,结合置信度阈值控制,在多轮解码下显著提升同声传译的质量-延迟权衡,并降低提交校准误差。
中文摘要 AI 辅助
同声传译必须在源语音完整之前输出有用的目标文本,同时保留每个已提交的标记。我们使用从模型自身的完整波形和部分波形翻译中导出的前缀监督来适配一个全话语语音语言模型,这既不需要转录文本,也不需要人工翻译。我们比较了单轮强制前缀解码和多轮仅追加解码,使用置信度阈值来控制推理时的质量-延迟权衡,并通过单独的合成余量来改变训练前缀的密度。在FLEURS和CoVoST2上,针对三种语言方向,前缀训练在质量-延迟前沿上优于未适配的模型,并且置信度提供了最广泛且持续具有竞争力的操作范围。多轮解码在低延迟下通常更强;在多轮训练下,提交校准误差总体下降63%-68%,在早期前缀处下降68%-80%,而单轮训练仅提供适度的总体校准改进,且没有早期前缀的改进。较小的合成余量有时会将前沿扩展到更低的延迟,特别是在较短的语音上,而较大的余量会降低翻译质量和校准效果。因此,前缀适配改善了同声传译,特别是在多轮仅追加解码下,而合成密度引入了非单调的质量-延迟权衡。
英文摘要
Simultaneous speech translation must emit useful target text before the source is complete while preserving every committed token. We adapt a full-utterance speech language model using prefix supervision derived from its own complete- and partial-waveform translations, requiring neither transcripts nor human translations. We compare single-turn forced-prefix and multi-turn append-only decoding, use a confidence threshold to control the inference-time quality--latency trade-off, and vary the density of training prefixes with a separate synthesis margin. On FLEURS and CoVoST2 in three language directions, prefix training improves quality--latency frontiers over the unadapted model, and confidence provides the broadest consistently competitive operating range. Multi-turn decoding is generally stronger at low latency; under multi-turn training, commit-calibration error falls by 63--68% overall and 68--80% at early prefixes, whereas single-turn training provides only modest overall calibration gains and no early-prefix improvement. A small synthesis margin sometimes extends the frontier to lower latency, particularly on shorter utterances, while a larger margin degrades translation quality and calibration. Prefix adaptation therefore improves simultaneous speech translation, especially under multi-turn append-only decoding, while synthesis density introduces a non-monotonic quality--latency trade-off.
发表机构
- Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。