arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一切皆在时机:面向基于大语言模型的同声语音到语音翻译的因果感知框架

All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation

Amir Hussein, Enas Albasiri, Travis M. Bartley, Nourchene Ferchichi, Ke Hu, Harishchandra Dubey, Myungjong Kim, Zhehuai Chen, Oluwatobi Olabiyi, Sanjeev Khudanpur

arXiv 2609.30416首次发表:更新:

发表机构

Johns Hopkins University; NVIDIA(约翰斯·霍普金斯大学; 英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大语言模型同声语音翻译中数据稀缺与延迟问题,提出因果感知框架FAST-CAP,通过分解架构与自适应策略,在CVSS上提升BLEU达1.2并降低延迟38.8%。

AI 中文摘要

大语言模型(LLMs)在低资源离线翻译中展现出强大性能;然而,由于缺乏具有高跨语言说话人保真度的因果对齐训练数据,将其扩展到同声语音到语音翻译(Simul-S2ST)仍然具有挑战性。此外,现有方法依赖于固定的翻译策略或置信度启发式方法,导致次优的质量和更高的延迟。我们提出了一种因果感知的Simul-S2ST框架,并配备了一种新颖的数据管道,该管道生成高保真、因果对齐的片段,并改善了语音迁移。该框架引入了(i)分解式S2ST架构(FAST),(ii)因果感知自适应策略(CAP),以及(iii)因果感知延迟度量。在CVSS西班牙语、德语和法语上的实验表明,FAST-CAP持续改善了质量-延迟权衡,与固定策略相比,实现了高达+1.2 BLEU的提升和26%的相对延迟降低。尽管使用的训练数据远少于现有系统,FAST-CAP在语音翻译质量和说话人保真度方面达到了最先进的结果,同时实现了高达38.8%的相对延迟降低。

英文摘要

Large Language Models (LLMs) have shown strong performance in low-resource offline translation; however, extending them to simultaneous speech-to-speech translation (Simul-S2ST) remains challenging due to the scarcity of causally aligned training data with high cross-lingual speaker fidelity. In addition, existing approaches rely on fixed translation policy or confidence heuristics, leading to suboptimal quality and higher latency. We propose a causality-aware Simul-S2ST framework with a novel data pipeline that generates high-fidelity, causally aligned segments with improved voice transfer. The framework introduces (i) a factorized S2ST architecture (FAST), (ii) a causality-aware adaptive policy (CAP), and (iii) causality-aware latency metric. Experiments on CVSS Spanish, German, and French show that FAST-CAP consistently improves the quality-latency trade-off, achieving up to +1.2 BLEU and a 26% relative latency reduction over a fixed policy. Despite using substantially less training data than existing systems, FAST-CAP achieves state-of-the-art results in speech translation quality and speaker fidelity while yielding up to a 38.8% relative reduction in latency.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑