arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于注意力的自适应策略用于同声语音到文本翻译

Attention-Based Adaptive Policies for Simultaneous Speech-to-Text Translation

Filip Tăşădan, Ema Tomanová, Ondrej Lopuch, Paweł Bilko, Anders Søgaard

arXiv 2609.30839首次发表:更新:

发表机构

University of Copenhagen(哥本哈根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出RFAP和DCAP两种基于注意力的自适应策略,使离线训练的语音翻译模型可用于流式场景,在CVSS-C上提升BLEU达4.0并减少延迟近1秒。

AI 中文摘要

同声语音到文本翻译(Simul-S2TT)包括在系统处理传入音频帧的同时生成部分翻译。然而,这种设置的流式特性带来了在最小化延迟的同时决定执行准确翻译的最佳时刻的挑战。为了解决这一挑战,我们利用编码器-解码器架构的交叉注意力机制来找到输入语音帧与目标文本标记之间的正确对齐。在本文中,我们提出了最近帧注意力策略(RFAP)和双条件注意力策略(DCAP),它们允许离线训练的语音到文本翻译模型在流式场景中使用,而无需额外训练。在CVSS-C语料库上的三种不同语言翻译对上的结果表明,RFAP能够超越其他策略,获得高达4.0 BLEU的提升,同时将翻译延迟减少近1秒。此外,DCAP能够在延迟非常低时保持较高的翻译质量。

英文摘要

Simultaneous speech-to-text translation (Simul-S2TT) consists of generating partial translations while the incoming audio frames are processed by the system. However, the streaming nature of this setup creates the challenge of deciding the best moment to perform an accurate translation while minimizing the delay. To address this challenge, we utilize the cross-attention mechanism of the encoder-decoder architecture to find the right alignment between the input speech frames and the target text tokens. In this paper, we propose the Recent Frame Attention Policy (RFAP) and the Dual-Condition Attention Policy (DCAP) that allow offline trained speech-to-text translation models to be used in streaming scenarios without requiring additional training. Results on three different language translation pairs over the CVSS-C corpus show that the RFAP is able to surpass other policies with gains of up to 4.0 BLEU while reducing the translation delay by almost 1 second. Moreover, the DCAP is able to preserve a high translation quality when the latency is very low.

CommentsSubmitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑