发表机构
Karlsruhe Institute of Technology; Texas A&M University; ETH Zurich; Carnegie Mellon University(卡尔斯鲁厄理工学院; 德克萨斯农工大学; 苏黎世联邦理工学院; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对语音翻译系统去除不流畅言语导致语义丢失的问题,构建了Uh-Mazing基准,发现错误起始和自我修正致翻译质量下降,提出推理时解码可缓解问题并发布相关资源。
AI 中文摘要
当前语音翻译系统(包括语音大模型SpeechLLMs)均基于清理后的文本进行训练,往往会去除填充停顿、错误起始等不流畅言语,而非对其进行翻译。我们表明这会带来代价:不流畅言语承载着被清理语音时丢失的语义。为系统研究该问题,我们推出Uh-Mazing基准,这是一个经人工翻译、标注了不流畅言语的Switchboard语音数据集,涵盖英语到八种目标语言的翻译任务。在这些语言及多种架构上,我们发现错误起始和自我修正(而非填充停顿或话语标记)是导致翻译质量下降的主要因素,且无法保留不流畅言语的模型往往会省略它而非错误翻译。我们还表明,推理时的解码无需重新训练即可缓解该问题,并发布了该基准及代码。
英文摘要
Current speech translation systems, including SpeechLLMs, are trained on cleaned text and tend to strip disfluencies like filled pauses and false starts rather than translate them. We show this comes at a cost: disfluencies carry meaning that gets lost when speech is cleaned up. To study this systematically, we introduce Uh-Mazing, a benchmark of human-translated, disfluency-annotated Switchboard speech covering English into eight target languages. Across these languages and several architectures, we find that false starts and self-repairs, not filled pauses or discourse markers, drive most of the translation-quality loss, and that models which fail to preserve a disfluency tend to omit it rather than mistranslate it. We show inference-time decoding can mitigate this without retraining, and release the benchmark and code.