arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2610.07545cs.CL

边缘设备上的质量感知自校正语音翻译

Quality-Aware Self-Correcting Speech Translation on an Edge Device

  • University of Surrey(萨里大学)
  • Institute for People-Centred AI(以人为本人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

Zubair Ajmal Farooq, Diptesh Kanojia

AI总结:

提出一种在 Jetson Nano 上运行的离线语音到语音翻译流水线,利用质量估计门控触发二次校正,发现 QE 作为门控有效但作为排序器不佳,移除它可释放内存且不损质量,并发布系统支持六语言对实时翻译。

AI中文摘要:

我们提出了一种完全离线的语音到语音翻译流水线,该流水线在 Jetson Nano(4 GB)上运行,无需重新训练即可纠正自身较弱的翻译。一个 Whisper-tiny 自动语音识别(ASR)模块将语音转换为文本,并馈送给 Opus-MT 翻译器;多语言 BERT 余弦相似度作为质量估计(QE)门控,当置信度低于预定义阈值 $\ au$ 时,触发二次通道校正。我们比较了三种校正方法:QE 重排序(M1)、最小贝叶斯风险解码(M2)和约束波束搜索(M3)。在 1,012 个 FLORES-200 句子(英语-西班牙语)上,M2 在 $\ au=0.90$ 时,与贪婪解码相比,在 BLEU(+0.67,p<0.001)、ChrF(+0.51,p<0.001)和 COMET(+0.0020,N=3,p=0.002)上产生了统计显著的改进;M1 未产生显著增益,而 M3 显著差于基线(p>0.99)。我们的核心发现是,QE 作为门控有效,但作为排序器效果不佳:从候选选择中移除 QE 模型(M1$\ o$M2)不会损害质量,并从关键路径中释放了 680 MB 内存。利用改编自后期编辑工作量文献的增益与编辑比率,我们进一步表明,较小的候选池(N=3)能产生更精准的校正,具有更好的语义充分性,而较大的候选池(N=10)能最大化词汇奖励。我们发布了该系统,并展示了跨六个语言对的实时翻译。

英文摘要:

We present a fully offline speech-to-speech translation pipeline that runs on a Jetson Nano (4 GB) and corrects its own weak translations without retraining. A Whisper-tiny ASR feeds an Opus-MT translator; multilingual BERT cosine similarity acts as a Quality Estimation (QE) gate, triggering a secondary-pass correction when confidence falls below a pre-defined threshold $τ$. We compare three correction methods: QE reranking (M1), Minimum Bayes-Risk decoding (M2), and constrained beam search (M3). On 1,012 FLORES-200 sentences (English-Spanish), M2 at $τ=0.90$ produces statistically significant improvements over greedy decoding on BLEU (+0.67, p<0.001), ChrF (+0.51, p<0.001), and COMET (+0.0020 at N=3, p=0.002); M1 yields no significant gains, and M3 is significantly worse than baseline (p>0.99). Our central finding is that QE functions effectively as a gate but poorly as a ranker: removing the QE model from candidate selection (M1$\to$M2) does not hurt quality and frees 680 MB from the critical path. Using a gain-to-edit ratio adapted from the post-editing-effort literature, we further show that smaller candidate pools (N=3) yield more surgical corrections with better semantic adequacy, while larger pools (N=10) maximise lexical reward. We release the system and demonstrate live translation across six language pairs.

补充信息

↑