AI 中文总结
针对机器翻译模型评估的饱和问题,提出对抗翻译优化(ATO)方法增强基准文本以提升翻译难度,生成了两个各含350篇英文文本的数据集,验证了其可有效降低翻译质量且无需额外成本。
AI 中文摘要
当最先进的机器翻译模型在标准基准上达到饱和时,该领域需要更具挑战性的评估来区分不同质量的模型。我们提出通过结合对抗优化与可微翻译难度估计器,对现有基准进行增强以提升翻译难度。我们的对抗翻译优化(Adversarial Translation Optimization,ATO)使用难度与流畅度联合目标的梯度,迭代替换token。由于每一步会在每个位置对候选替换进行分支,优化变为树搜索问题,我们采用束搜索(Beam Search)解决。ATO提供了一种基于梯度的替代方案,无需LLM提示、昂贵的人工整理或特定任务模型训练即可创建数据集。我们经ATO修改的基准将平均翻译质量(xCOMET)从0.93降至0.82,相比之下, paraphrasing方法为0.88,零样本基线为0.86。人工评估显示,修改后的文本虽比基线稍欠自然,但仍保持合理的语法性与合理性,同时翻译难度显著提升。我们发布了通过该方法生成的各含350篇英文文本的两个数据集及代码。
英文摘要
As state-of-the-art machine translation models saturate standard benchmarks, the field needs more challenging evaluations to distinguish between models of varying quality. We propose augmenting existing benchmarks to increase translation difficulty by combining adversarial optimization with a differentiable translation difficulty estimator. Our Adversarial Translation Optimization (ATO) uses gradients from a combined difficulty and fluency objective to iteratively replace tokens. Because each step branches over candidate substitutions at every position, optimization becomes a tree search problem, which we address with Beam Search. ATO offers a gradient-based alternative to LLM-based dataset creation without LLM prompting, expensive human curation, or task-specific model training. Our ATO-modified benchmark lowers average translation quality (xCOMET) from 0.93 to 0.82, compared to 0.88 for paraphrasing and 0.86 for a zero-shot baseline. Human evaluation shows the modified texts are somewhat less natural than the baselines but remain reasonably grammatical and plausible while being substantially harder to translate. We release two datasets of 350 English texts each, generated by our methods, as well as the code.
Comments18 pages, 8 figures, 10 tables. William Kalikman and Šimon Sukup contributed equally. Published in EAMT 2026. Code: https://github.com/BreakingMT/ATO. Data: https://huggingface.co/datasets/wskal/ATO-datasets