arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.12674cs.CL

Doc2FRC:通过固定范围分块实现长度一致的文档级机器翻译

Doc2FRC: Length-Consistent Document-Level Machine Translation via Fixed-Range Chunking

  • The University of Tokyo(东京大学)
  • Kyoto University(京都大学)
  • Riken(理化学研究所)
  • Tohoku University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

Xiaotian Wang, Youyuan Lin, Zhan Shen, Hitomi Yanaka

AI总结:

本文提出固定范围分块(FRC)方法,通过动态规划将文档划分为长度一致的块,减少训练与推理长度不匹配,并配合双边界匹配算法和多种训练策略,显著提升7B大语言模型的文档级翻译质量,在IWSLT2017及新构建的10语言测试集上均优于现有方法。

AI中文摘要:

具有长上下文窗口的高级大语言模型(LLMs)可以大幅减少文档级机器翻译(DocMT)中的输入截断问题。然而,直接的Doc2Doc翻译仍然容易出现n-gram重复和渐进式质量下降。一种常见的补救措施是将文档分割成更细粒度的块。尽管如此,传统的基于规则的分块方法无法处理训练与推理之间的长度分布不匹配问题。为了解决这个问题,我们引入了固定范围分块(FRC),利用动态规划将文档划分为预定义长度区间内的块。通过在训练和推理过程中一致地应用FRC,任意长度的输入文档都被映射到相同的长度分布,从而大幅减少训练与测试之间的长度不匹配。以FRC为核心,我们提出了一种轻量级的双边界匹配算法用于块对齐,以及四种不同的训练策略。实验结果表明,基于FRC的微调在7B大语言模型上相比直接Doc2Doc微调有显著提升,并在IWSLT2017上优于现有的DocMT方法。我们进一步构建了GlobVDoc,一个独立于主流DocMT训练来源的10语言测试集,并表明FRC改善了分布外文档翻译的质量。

英文摘要:

Advanced large language models (LLMs) with long context windows can substantially reduce input truncation in document-level machine translation (DocMT). However, direct Doc2Doc translation remains prone to n-gram repetition and progressive quality degradation. A common remedy is to segment the document into finer-grained chunks. Nonetheless, conventional rule-based chunking approaches fail to handle the length distribution mismatch between training and inference. To address this, we introduce Fixed-Range Chunking (FRC), utilizing dynamic programming to partition documents into chunks within a predefined length interval. By consistently applying FRC during training and inference, the input documents of any length are mapped to the same length distribution, substantially reducing train-test length mismatch. Centered on FRC, we propose a lightweight dual-boundary matching algorithm for chunk alignment, alongside four distinct training strategies. Experimental results show that FRC-based fine-tuning substantially improves 7B LLMs over direct Doc2Doc fine-tuning and outperforms existing DocMT methods on IWSLT2017. We further construct GlobVDoc, a 10-language test set independent of mainstream DocMT training sources, and show that FRC improves out-of-distribution document translation.

补充信息

↑