发表机构
University of Michigan; Peking University; Microsoft Research; Northwestern University(密歇根大学; 北京大学; 微软研究院; 西北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对文本评估指标,提出统计与策略对齐概念及三项测试原则,开发基于互信息的指标设计框架,发现 LLM-as-a-Judge 相关性高但易操纵,新指标鲁棒性强且相关性具竞争力。
AI 中文摘要
基于参考的文本评估指标被广泛用于评估自然语言生成系统,通过将候选响应与参考响应比较来对候选响应打分。评估指标的可靠性通常由其与人类评分的统计相关性判断,但随着这些指标越来越多地被用作优化目标,仅靠相关性已不足够:智能体可能会策略性地操纵评估指标。我们通过两个互补的对齐概念研究这一问题:若指标与人类评分相关则为统计对齐,若指标能抵抗未添加任务相关信息的扰动则为策略对齐。我们有两项贡献:第一,提出了基于参考的指标的测试原则,包括人类评分相关性、退化敏感性和操纵鲁棒性,这些原则用于评估指标是否符合人类判断、惩罚低 effort 的信息损失、抵抗策略性分数膨胀;第二,开发了基于互信息的指标的统一设计框架,该框架将现有及新指标分解为信息度量、估计方法、文本表示和预测机制四个选择。在同行评审、摘要生成和问答任务中,我们发现高人类评分相关性并不意味着策略对齐:LLM-as-a-Judge 实现了高相关性但易受操纵,而基于互信息的指标显著提升了操纵鲁棒性。我们的框架还发现了一个新指标,在实验中实现了最强的整体鲁棒性,同时在人类评分相关性上仍具有竞争力。
英文摘要
Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with a reference response. The reliability of an evaluation metric is usually judged by its statistical correlation with human ratings. However, as these metrics are increasingly used as optimization objectives, correlation alone is no longer sufficient: agents may strategically game the evaluation metric. We study this issue through two complementary notions of alignment. A metric is statistically aligned if it correlates with human ratings and strategically aligned if it resists perturbations that do not add task-relevant information. We make two contributions. First, we propose test principles for reference-based metrics consisting of human-rating correlation, degradation sensitivity, and manipulation robustness. These principles evaluate whether a metric agrees with human judgments, penalizes low-effort information loss, and resists strategic score inflation. Second, we develop a unified design framework for mutual-information-based metrics that decomposes existing and new metrics into four choices: information measure, estimation method, text representation, and prediction mechanism. Across peer review, summarization, and question answering, we find that strong human-rating correlation does not imply strategic alignment: LLM-as-a-Judge achieves high correlation but is susceptible to manipulation. In contrast, mutual-information-based metrics substantially improve manipulation robustness. Our framework also uncovers a new metric that achieves the strongest overall robustness in our experiments while remaining competitive on human-rating correlation.