通过GRPO生成机器翻译对抗文本
Generating Adversarial Texts for Machine Translation via GRPO
浏览论文内容
中文总结 AI 辅助
提出基于GRPO强化学习的方法,通过微调大语言模型重写源文本,生成更难翻译的对抗样本,显著降低翻译质量并保持可读性。
中文摘要 AI 辅助
随着机器翻译(MT)系统的不断改进,标准基准在揭示其剩余弱点方面变得不那么有效。传统上,创建具有挑战性的测试集的方法依赖于昂贵的人工创建或筛选,而自动化方法则难以生成具有必要翻译难度和语言多样性的测试集。我们提出了一种基于强化学习的可扩展方法,用于将现有源文本改写为对MT系统而言更难翻译的实例。我们使用组相对策略优化(GRPO)微调大型语言模型,采用基于翻译难度的奖励信号,并结合语义相似性、语法正确性和近似长度保持的约束。在WMT25上,我们的方法将平均COMET翻译质量从0.63大幅降低至0.48,同时保持了语法正确性和可读性,而基础模型仍保持在0.64。对未见过的WMT19-WMT24基准的评估证实,这种行为能够泛化到训练数据之外,人工评估进一步表明,改写后的文本显著降低了翻译质量,同时自然度适度下降,语法正确性仅有微小变化。我们公开代码以支持可复现性。
英文摘要
As machine translation (MT) systems continue to improve, standard benchmarks become less informative for exposing remaining weaknesses. Traditional methods for creating challenging test sets rely on expensive manual creation or curation, while automated approaches struggle to produce sets with the necessary translation difficulty and linguistic diversity. We propose a scalable reinforcement-learning-based approach for rewriting existing source texts into instances that are more difficult to translate for MT systems. We fine-tune a large language model with Group Relative Policy Optimization (GRPO), using reward signals based on translation difficulty together with constraints for semantic similarity, grammaticality, and approximate length preservation. On WMT25, our approach substantially reduces average COMET translation quality from 0.63 to 0.48, while preserving grammaticality and readability, whereas the base model remains at 0.64. Evaluations on the unseen WMT19-WMT24 benchmarks confirm that this behavior generalizes beyond the training data, and human evaluation further shows that the rewrites substantially lower translation quality while incurring a moderate drop in naturalness and only a small change in grammaticality. We release our code to support reproducibility.
发表机构
- ETH Zurich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。