Position: Evaluation Scores Are Perishable Knowledge Claims
Position: 评估分数是易逝的知识主张
机构 * DeepThought Solutions(深度思考解决方案公司) ; Meta ; University of Florida(佛罗里达大学)
AI总结 该研究指出语言模型评估存在信任膨胀问题,提出将评估分数视为具有形式性、范围性和有效性窗口的认知主张,建议添加元数据并采用最弱链聚合,发现HELM排行榜上两种聚合方式的前五名模型完全不重合。
Comments 7 pages, 1 figure, 1 table. Published at the Fifth Workshop on Generation, Evaluation and Metrics (GEM), ACL 2026, San Diego
Journal ref Proceedings of the Fifth Workshop on Generation, Evaluation and Metrics (GEM), ACL 2026, pages 1029-1035