PAMT:面向多领域机器翻译的过程对齐强化学习
PAMT: Process-Aligned Reinforcement Learning for Multi-Domain Machine Translation
浏览论文内容
中文总结 AI 辅助
针对多领域机器翻译中显式推理的偏差问题,提出过程对齐强化学习框架PAMT,结合长思维链监督与多维度奖励机制,在多种场景下表现优于基线模型。
中文摘要 AI 辅助
多领域机器翻译(MDMT)不仅需要流畅的生成,还要求做出领域敏感的翻译决策,如领域消歧、术语控制和风格适配。大推理模型(LRM)通过中间翻译步骤明确做出此类决策,但我们对15个领域和4种翻译方向的分析显示,这种显式推理是双刃剑:它提升了长文本和高难度翻译的质量,却常在术语密集、风格受限的场景中出现偏差。我们将该失效归因于信用分配瓶颈:现有方法仅优化最终输出或粗粒度轨迹,无法识别哪些翻译步骤真正对最终翻译有帮助。为解决此问题,我们提出PAMT,即结合冷启动领域感知长思维链(Long-CoT)监督与强化学习的过程对齐训练框架。PAMT针对最终翻译使用序列级格式奖励和结果奖励,同时使用步骤级过程奖励,衡量每个显式翻译步骤提升参考翻译概率的程度。在两种主干模型上,PAMT均优于基础模型,平均表现优于机器翻译(MT)专用基线,且在领域内、分布外(OOD)和多语言场景中与强大语言模型(LLM)/大推理模型(LRM)具有竞争力。
英文摘要
Multi-domain machine translation (MDMT) requires more than fluent generation: it demands domain-sensitive translation decisions such as domain disambiguation, terminology control, and stylistic adaptation. Large reasoning models (LRMs) make such decisions explicit through intermediate translation steps, but our analysis across 15 domains and four translation directions shows that this explicit reasoning is double-edged: it improves long-form and high-difficulty translation, yet often drifts in terminology-intensive and stylistically constrained settings. We trace this failure to a credit-assignment bottleneck: existing methods optimize final outputs or coarse trajectories, but cannot identify which translation steps actually help the final translation. To address this, we propose PAMT, a process-aligned training framework that combines cold-start domain-aware Long-CoT supervision with reinforcement learning. PAMT uses sequence-level format and outcome rewards for the final translation, together with a step-level process reward that measures how much each explicit translation step increases the likelihood of the reference translation. Across two backbones, PAMT improves over base models, outperforms MT-specialized baselines on average, and remains competitive with strong LLMs/LRMs across in-domain, OOD, and multilingual settings.