arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StalePO:利用机器翻译中的旧版后编辑进行锚定令牌级偏好优化

StalePO: Anchored Token-Level Preference Optimization using Legacy Post-Edits in Machine Translation

Rohit Dhaipule, Sukhdeep Singh Kharbanda, Prasanth Bathala, Pradyumna Lanka, Anubhav Shrimal

arXiv 2609.16340首次发表:更新:

发表机构

Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对机器翻译中旧版后编辑作为偏好信号的陈旧偏好问题,提出StalePO目标函数,通过双向似然下移、策略锚定和令牌级KL约束,在本地化数据上显著提升MQM质量检查通过率。

AI 中文摘要

机器翻译系统会定期升级为更强的模型,但可用的偏好信号是旧系统输出的人工后编辑,而新模型可能已经超越了这些输出。此外,为每个新模型收集新的后编辑成本过高,令人望而却步。我们将此称为“陈旧偏好”问题。标准DPO在此设置下可能失效:它可能增加低质后编辑的似然性,侵蚀模型现有质量,并且无法提供纠正局部错误所需的逐令牌控制。我们引入StalePO,这是一个源于该场景三个要求的目标函数。似然移动必须在两个响应上都向下,策略必须锚定到其自身的基础响应,且KL约束必须应用于令牌级别。这些要求是共同必要的。在消融实验中,每个机制单独使用都使模型性能与基础模型无显著差异,只有它们的组合才能将陈旧反馈转化为收益。在英语到印地语和英语到土耳其语的本地化数据上,StalePO将通过所有LLM-as-judge MQM质量检查的片段比例分别提高了14.9和4.6个百分点,收益集中在风格和流畅性上。在同一框架下进行的人工评估证实了这些在英语到印地语上的收益,将通过所有七项人工检查的片段比例提高了13.8个百分点。

英文摘要

Machine translation systems are periodically upgraded to stronger models, but the available preference signal is human post-edits of an older system's outputs, which the newer model may already surpass. Moreover, collecting fresh post-edits for every new model is prohibitively expensive. We call this the Stale Preference problem. Standard DPO can fail in this setting: it may increase the likelihood of inferior post-edits, erode the model's existing quality, and fail to provide the per-token control needed to correct localized errors. We introduce StalePO, an objective derived from three requirements this regime imposes. Likelihood movement must be downward on both responses, the policy must be anchored to its own base response, and the KL constraint must apply at the token level. These requirements are jointly necessary. In ablations, each mechanism in isolation leaves the model's performance indistinguishable from the base model, and only their combination converts stale feedback into gains. On English-to-Hindi and English-to-Turkish localization data, StalePO improves the fraction of segments passing all LLM-as-judge MQM quality checks by 14.9 and 4.6 percentage points, respectively, with gains concentrated on style and fluency. A human evaluation under the same framework confirms these gains on English-to-Hindi, raising the fraction of segments passing all seven human checks by 13.8 percentage points.

CommentsAccepted to the 11th Conference on Machine Translation (WMT 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑