arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

X-Translator:一种实时多语言说话人感知语音到语音翻译系统

X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System

Yuxiang Zhao, Yichi Zhang, Yanjie An, Yanqiao Zhu, Zhanxun Liu, Yushen Chen, Qixi Zheng, Haina Zhu, Yunchong Xiao, Keqi Deng, Shuai Fan, Kai Yu, Xie Chen

arXiv 2607.17544首次发表:更新:

发表机构

School of Computer Science, Shanghai Jiao Tong University; Shanghai Innovation Institute; MoE Key Lab of Artificial Intelligence; Jiangsu Key Lab of Language Computing(上海交通大学计算机科学系; 上海创新研究院; 教育部人工智能重点实验室; 江苏省语言计算重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在实现实时多语言语音到语音翻译,提出低成本模块化级联系统X-Translator,结合流式ASR、机器翻译和提示条件TTS,经实验评估翻译等指标,为理解面向部署的S2ST实际权衡提供开放平台。

AI 中文摘要

实时语音到语音翻译(S2ST)系统必须在翻译质量、延迟、语音自然度和说话人一致性之间取得平衡。公开记录的S2ST系统在直接、多语言、流式和表达性建模方面取得了进展,而专有产品和API越来越多地向用户展示实时翻译功能。然而,对于开放和可复制的系统来说,实际部署仍然具有挑战性,特别是在长格式和多说话人对话中,部分ASR假设不稳定,轮次边界模糊,并且目标语音必须通过适当的说话人提示生成。我们提出了X-Translator,这是一种低成本的模块化级联S2ST系统,它通过会话级运行时控制器结合了流式ASR、机器翻译和提示条件TTS。该系统使用增量段承诺将不稳定的ASR流转换为可翻译的单元,并使用在线说话人提示管理器将源语音跨度绑定到特定说话人的语音提示以进行合成。我们使用OpenSTBench评估翻译、语音质量和延迟,与专有语音翻译API作为行为基线进行比较,测量长格式语音稳定性,评估多说话人对话中的说话人保留情况,并评估多语言翻译质量。X-Translator提供了一个开放平台,用于理解面向部署的S2ST的实际权衡。代码和演示可在此https URL获得。

英文摘要

Real-time speech-to-speech translation (S2ST) systems must balance translation quality, latency, speech naturalness, and speaker consistency. Publicly documented S2ST systems have advanced direct, multilingual, streaming, and expressive modeling, while proprietary products and APIs increasingly expose real-time translation capabilities to users. However, practical deployment remains challenging for open and reproducible systems, especially in long-form and multi-speaker conversations where partial ASR hypotheses are unstable, turn boundaries are ambiguous, and target speech must be generated with an appropriate speaker prompt. We present X-Translator, a low-cost modular cascaded S2ST system that combines streaming ASR, machine translation, and prompt-conditioned TTS through a session-level runtime controller. The system uses incremental segment commitment to convert unstable ASR streams into translation-ready units, and an online speaker prompt manager to bind source speech spans to speaker-specific voice prompts for synthesis. We evaluate translation, speech quality, and latency with OpenSTBench, compare against proprietary speech translation APIs as behavioral baselines, measure long-form voice stability, evaluate speaker preservation in multi-speaker conversations, and assess multilingual translation quality. X-Translator provides an open platform for understanding the practical trade-offs of deployment-oriented S2ST. Code and demo are available at https://github.com/zhaoyx239/X-Translator.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑