超越聚合分数:评估基于参考的自动评估方法的行为正确性假设
Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods
- Oak Ridge National Laboratory(橡树岭国家实验室)
- DHS Science and Technology Directorate(美国国土安全部科学与技术局)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出行为正确性假设框架,评估多种类型的基于参考的自动评估方法,发现无评估器满足所有正确性假设,聚合性能相似的评估器行为特征差异显著,可提供传统元评估掩盖的诊断信息。
AI中文摘要:
基于参考的自动评估方法在评估自然语言生成系统中发挥关键作用。现有元评估主要衡量与人工判断或基准标签的一致性,对受控条件下评估器行为的洞察有限。我们引入行为正确性假设,这是用于评估基于参考的自动评估方法的补充框架。我们定义了保持正确性和改变正确性的假设分类,并通过受控响应转换将其实现,这些转换规定了预期的评分行为。我们评估了多种词汇级、字符级、语义级、基于LLM的及混合评估器,并分析它们在假设层面的行为、稳定性、敏感性、重复运行变异性、配置敏感性和可复现性。实验揭示了不同评估范式间存在明显的行为权衡:没有评估器满足所有提出的正确性假设,且聚合性能相似的评估器可呈现截然不同的行为特征。这些发现表明,行为正确性假设提供了被传统聚合元评估所掩盖的诊断信息。
英文摘要:
Automated reference-based evaluation methods play a critical role in assessing natural language generation systems. Existing meta-evaluation primarily measures agreement with human judgments or benchmark labels, providing limited insight into evaluator behavior under controlled conditions. We introduce behavioral correctness assumptions, a complementary framework for evaluating reference-based automatic evaluation methods. We define a taxonomy of correctness-preserving and correctness-altering assumptions and operationalize them through controlled response transformations that specify expected scoring behaviors. We evaluate diverse lexical, character-level, semantic, LLM-based, and hybrid evaluators and analyze their assumption-level behavior, stability, sensitivity, repeat-run variability, configuration sensitivity, and reproducibility. Our experiments reveal distinct behavioral trade-offs across evaluation paradigms: no evaluator satisfies all proposed correctness assumptions, and evaluators with similar aggregate performance can exhibit substantially different behavioral profiles. These findings demonstrate that behavioral correctness assumptions provide diagnostic information obscured by conventional aggregate meta-evaluation.