诊断,然后修复:面向领域特定机器翻译的两阶段MQM引导的后编辑框架
Diagnose, Then Repair: A Two-Stage MQM-Guided Post-Editing Framework for Domain-Specific Machine Translation
- Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出两阶段评估器引导的后编辑框架,将MQM诊断转化为最小修复,提升可控性并减少释义漂移,在七种语言和七个LLM上优于一阶段方法。
AI中文摘要:
基于大语言模型的机器翻译评估可以紧密匹配人类判断,但在实践中,它仍然主要是诊断性的,这些信号很少在真实生产约束下转化为直接的质量改进。我们提出了一个两阶段、评估器引导的自动后编辑框架,将MQM风格的评估转化为有针对性的修复:一个检索增强的LLM评估器在明确的编辑契约下输出结构化的、跨度级别的MQM诊断,而一个单独的LLM后编辑器仅对这些诊断应用最小编辑。这种分离相比一阶段“判断并细化”基线提高了可控性并减少了释义漂移。在一项涉及七个LLM、跨越三个模型提供商和七种语言的系统研究中,我们最佳配置在COMET-22和COMETKiwi分数上均一致优于一阶段后编辑方法,同时评估器的错误跨度和严重性与人类MQM标注及人类编辑偏好表现出强烈一致性。
英文摘要:
LLM-based machine translation evaluation can closely match human judgments, but in practice it remains largely diagnostic, with the signals rarely translating into direct quality improvements under real production constraints. We propose a two-stage, evaluator-guided automatic post-editing framework that turns MQM-style evaluation into targeted repairs: a retrieval-augmented LLM evaluator outputs structured, span-level MQM diagnoses under an explicit edit contract, and a separate LLM post-editor applies minimal edits restricted to those diagnoses. This separation improves controllability and reduces paraphrastic drift compared to one-stage "judge-and-refine" baselines. In a systematic study involving seven LLMs spanning three model providers and seven languages, our best configuration consistently improves both COMET-22 and COMETKiwi scores over one-stage post-edit methods, while the evaluator's error spans and severities show strong agreement with human MQM annotations and human editor preferences.