arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

适可而止:机器翻译重排序中高效候选生成的弃权策略

Quit While You're Ahead: Quit for Efficient Candidate Generation in Machine Translation Reranking

Guangyu Chen, Boxuan Lyu, Hidetaka Kamigaito, Kotaro Funakoshi, Manabu Okumura

arXiv 2609.00588首次发表:更新:

发表机构

Institute of Science Tokyo; Nara Institute of Science and Technology(东京科学学院; 奈良科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对机器翻译重排序的高延迟问题,提出Quit早停策略,在19种语言对的3个NMT模型上实现MBR和QE重排序的显著端到端加速,同时保持翻译质量。

AI 中文摘要

重排序方法,如最小贝叶斯风险(MBR)解码和质量估计(QE)重排序,在现代神经机器翻译(NMT)中被广泛用于从一组候选假设中选择输出。然而,性能提升是以高推理延迟为代价的。现有的加速方法仅针对MBR解码,仅减少重排序计算,未解决QE重排序问题,且候选生成这一可能更大的计算瓶颈基本未被触及。本研究提出Quit(量化不确定性以实现增量终止),一种针对整个生成-重排序流水线的新型早停策略。将候选生成视为不确定性下的序贯决策,Quit增量生成并重排序候选,当候选集中最高估计质量稳定时停止。在19种语言对的3个NMT模型上开展的综合实验表明,Quit为MBR带来1.47至2.66倍的端到端加速,为QE重排序带来3.43至4.12倍的加速,同时在预设的等价范围内保持翻译质量。

英文摘要

Reranking methods, such as Minimum Bayes Risk (MBR) decoding and Quality Estimation (QE) reranking, have been widely used in modern neural machine translation (NMT) to select an output from a set of candidate hypotheses. However, the performance gains come at the cost of high inference latency. Existing acceleration methods target MBR decoding and reduce only the reranking computation, leaving QE reranking unaddressed and candidate generation---which can be the larger computational bottleneck---largely untouched. In this work, we propose Quit (Quantifying Uncertainty for Incremental Termination), a novel early-stopping strategy for the entire generation--reranking pipeline. Quit treats candidate generation as a sequential decision-making process under uncertainty. It incrementally generates and reranks candidates, stopping when the best reranking score stabilizes. Comprehensive experiments with three NMT models across 19 language pairs show that Quit achieves end-to-end speedups of $1.47$--$2.66\times$ for MBR decoding and $3.43$--$4.12\times$ for QE reranking while preserving automatic metric scores.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑