发表机构
EURECOM(EURECOM)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出元评估框架,通过扰动金标准答案测试事实性评估指标,发现流水线方法优于LLM裁判,并提出成本效益高的新指标变体。
AI 中文摘要
评估大型语言模型(LLM)的事实正确性对许多应用至关重要。但我们的评估工具本身是否可信?尽管基于事实性的指标不断涌现,其敏感性和可靠性仍未得到充分探索。本文引入了一个元评估框架,通过受控的金标准答案扰动来系统测试这些指标。我们的方法生成具有已知退化程度的排序输出,以探究指标如何捕捉真实性的细微变化。实验表明,基于流水线的方法(如RAGAS的事实正确性指标)比LLM作为裁判的方法更能跟踪退化。我们还提出了一种新的事实正确性指标变体,该变体具有竞争力和成本效益。
英文摘要
Evaluating the factual correctness of large language models (LLMs) is vital for many applications. But are our evaluation tools themselves trustworthy? Despite the rise of factuality-based metrics, their sensitivity and reliability remain underexplored. This paper introduces a meta-evaluation framework that systematically tests these metrics using controlled corruptions of gold standard answers. Our method generates ranked outputs with known degrees of degradation to probe how metrics capture nuanced changes in truthfulness. Our experiments reveal that pipeline-based methods, such as the RAGAS's factual correctness metric, better track degradation than LLM-as-judge approaches. We also propose a new variant of the factual correctness metric that provides a competitive and cost-efficient.
Comments22 pages, 7 figures. Extended version of a paper accepted at EvalLLM 2025 (CORIA-TALN 2025)