超越基于参考的评估:用于语法纠错元评估的奖励模型
Beyond Reference-Based Evaluation: Reward Models for Meta-Evaluation of Grammatical Error Correction
- University of Waterloo(滑铁卢大学)
- Vector Institute(向量研究所)
- Scribendi Inc.(Scribendi公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出RM-EVAL,一种基于SEEDA人类偏好数据训练的无参考奖励模型,用于语法纠错的元评估,并通过奖励引导的文本生成(RGTG)改进GEC系统,实验证明其与人类判断高度一致且能提升生成质量。
AI中文摘要:
基于参考的语法纠错(GEC)评估指标(如M$^2$和ERRANT)假设参考集枚举了所有有效的编辑,因此常常会惩罚那些语法正确、保持语义但措辞不同的修正。我们引入了RM-EVAL,一个在SEEDA的人类偏好数据上训练的奖励模型,作为一个无参考的元评估器,能够在全序列和部分序列层面预测类似人类的质量判断。除了评估之外,我们展示了同一个奖励模型可以用作学习信号,通过奖励引导的文本生成(RGTG)来改进GEC生成,该方法保持基础GEC模型冻结,并执行在线、奖励驱动的解码。在SEEDA上,RM-EVAL与人类排名实现了强一致性,并且RGTG在奖励和外部验证方面取得了一致的提升,展示了一个无需依赖黄金参考即可评估和增强GEC系统的统一框架。
英文摘要:
Reference-based metrics for Grammatical Error Correction (GEC) such as M$^2$ and ERRANT assume that the reference set enumerates all valid edits, and therefore often penalize corrections that are grammatical and meaning-preserving but phrased differently. We introduce RM-EVAL, a reward model trained on human preference data from SEEDA, as a reference-free meta-evaluator that predicts human-like quality judgments at both full-sequence and partial-sequence levels. Beyond evaluation, we show that the same reward model can be used as a learning signal to improve GEC generation via Reward-Guided Text Generation (RGTG), which keeps a base GEC model frozen and performs online, reward-driven decoding. Across SEEDA, RM-EVAL achieves strong agreement with human rankings, and RGTG yields consistent gains in reward and external validation, demonstrating a unified framework for both assessing and enhancing GEC systems without relying on gold references.