arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Position: 评估分数是易逝的知识主张

Position: Evaluation Scores Are Perishable Knowledge Claims

Sankalp Gilda, Shlok Gilda

arXiv 2607.26191首次发表:更新:

发表机构

DeepThought Solutions; Meta; University of Florida(深度思考解决方案公司; Meta; 佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究指出语言模型评估存在信任膨胀问题,提出将评估分数视为具有形式性、范围性和有效性窗口的认知主张,建议添加元数据并采用最弱链聚合,发现HELM排行榜上两种聚合方式的前五名模型完全不重合。

AI 中文摘要

语言模型的评估方法正日益融合多种信号,涵盖自动指标、大语言模型(LLM)作为评判者的评分、人工评估以及基准套件结果。当这些信号通过平均法聚合时,评估置信度会大幅超过最弱信号的可靠性,我们将这一现象称为评估中的信任膨胀。我们认为,评估分数应被视为具有三个属性的认知主张:形式性(人工评估比自动指标提供更强证据)、范围性(基准结果仅适用于测试分布,而非通用)以及有效性窗口(随着污染累积和分布偏移,基准结果会失效)。多个趋同的研究传统(思维链分析、可能性逻辑和代数理论)确立了最弱链聚合是由单一悲观参数控制的参数化算子族的保守端点。基于这些传统,以及构建智能体AI评估工具的具体经验,我们提出评估结果应附带明确元数据(形式性层级、范围声明和到期日期),以使其认知状态透明。我们在公开的HELM排行榜上展示了平均聚合的代价:在十个场景下的54个前沿模型中,按平均分数和最弱链排名的前五名模型完全不重合。

英文摘要

Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessments and benchmark suite results. When these signals are aggregated via averaging, evaluation confidence can then substantially exceed the reliability of the weakest signal: a phenomenon we call trust inflation in evaluation. We argue that evaluation scores should be treated as epistemic claims with three properties: formality (human evaluation provides stronger evidence than an automated metric), scope (a benchmark result applies to the tested distribution, not universally), and validity windows (benchmark results expire as contamination accumulates and distributions shift). Several converging research traditions (chain-of-thought analysis, possibilistic logic, and algebraic theory) establish weakest-link aggregation as the conservative endpoint of a parameterized operator family controlled by a single pessimism parameter. Drawing on those traditions, and on concrete lessons from building an evaluation harness for agentic AI, we propose that evaluation results carry explicit metadata (formality tier, scope declaration, and expiration date) to make their epistemic status transparent. We illustrate the cost of mean aggregation on the public HELM leaderboard: across 54 frontier models on ten scenarios, the top-five models ranked by mean score and by weakest-link are completely disjoint.

Comments7 pages, 1 figure, 1 table. Published at the Fifth Workshop on Generation, Evaluation and Metrics (GEM), ACL 2026, San Diego

Journal refProceedings of the Fifth Workshop on Generation, Evaluation and Metrics (GEM), ACL 2026, pages 1029-1035

DOI:10.18653/v1/2026.gem-main.80

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑