不要数编辑次数,仅以结果评判:基于奖励的语法纠错评估
Don't Count the Edits, Judge by the Outcome Alone: Reward-Based Evaluation for Grammatical Error Correction
- Chung-Ang University(中央大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出SURE,一种基于源条件奖励的语法纠错评估器,通过联合学习整体奖励与多标准监督,在SEEDA上优于基线,尤其擅长改写式修正。
AI中文摘要:
语法纠错(GEC)的评估传统上依赖于参考或编辑重叠,这可能会惩罚与黄金修正不同的有效改写。无参考指标减少了这种依赖,但评估流畅的输出是否为源句的有效修正仍然具有挑战性。我们提出了SURE,一种源条件奖励评估器,在源内偏好上训练,涵盖最小编辑和改写导向的修正。SURE联合学习整体奖励与语法性、忠实性和流畅性的标准级监督,以及源侧错误解决的跨度级基础。在SEEDA上的实验表明,SURE与强基线相比表现具有竞争力,尤其在改写式修正和更解耦的标准级诊断上有显著提升。我们的代码可在以下网址获取:此https URL。
英文摘要:
Grammatical error correction (GEC) evaluation has traditionally relied on reference or edit overlap, which can penalize valid rewrites that differ from gold corrections. Reference-free metrics reduce this dependence, but evaluating whether a fluent output is a valid correction of the source remains challenging. We propose SURE, a source-conditioned reward evaluator trained on within-source preferences spanning minimal-edit and rewrite-oriented corrections. SURE jointly learns an overall reward with criteria-level supervision for grammaticality, faithfulness, and fluency, together with span-level grounding for source-side error resolution. Experiments on SEEDA show that SURE performs competitively against strong baselines, with particular gains on rewrite-style corrections and more disentangled criteria-level diagnostics. Our code is available at https://github.com/hayeonggg/SURE.