arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不要相信基准测试:通用LLM排名的局限性及任务特定评估的案例

Do Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific Evaluation

Danial Amin

arXiv 2609.23201首次发表:更新:

AI 中文总结

本文指出通用LLM排名存在五类局限性,主张采用披露配置、验证问题、报告成本及明确泛化范围的任务特定评估,并以众包平台Isotanta为例说明,模型选择应基于预期任务的性能证据而非排行榜高分。

AI 中文摘要

基准测试分数日益影响着大型语言模型(LLM)的开发、营销和选择。然而,一个总体分数只有在与所测试的系统、所包含的问题以及评估条件相关时才具有可解释性。本文审视了通用LLM排名的五个相互关联的局限性:被评估系统与公开可用系统之间的差异;外部评估中的商业激励和依赖;基准饱和、有缺陷的测试和数据污染;模型利用评分程序的行为;以及通用分数对用户任务的有限相关性。有记录的案例说明了为什么这些问题需要不同的应对措施。我主张评估程序应披露所测试的配置,验证问题和成功的任务完成情况,报告性能以及成本和执行时间,并明确泛化的范围。然后,我讨论了Isotanta,一个众包的基准测试平台,作为贡献问题和重复评估的实际例子。更大的问题池可能改善任务覆盖,而重复采样可以提高该池上估计的稳定性;两者都不能保证有效性或个性化。本文区分了该平台当前的共享排名与提出的任务特定和用户提供的评估。其核心论点是,模型选择需要关于预期工作性能的证据,而不仅仅是通用排行榜上的高排名。

英文摘要

Benchmark scores increasingly influence the development, marketing, and selection of large language models (LLMs). Yet an overall score is interpretable only in relation to the system tested, the questions included, and the conditions of evaluation. This perspective examines five connected limitations of general LLM rankings: differences between evaluated and publicly available systems; commercial incentives and dependencies in external evaluation; benchmark saturation, defective tests, and data contamination; models exploiting scoring procedures; and the limited relevance of general scores to users' tasks. Documented cases illustrate why these problems require different responses. I argue for evaluation procedures that disclose the tested configuration, validate questions and successful task completion, report performance alongside cost and execution time, and make the scope of generalization explicit. I then discuss \textbf{Isotanta}, a crowdsourced benchmarking platform, as a practical example of contributed questions and repeated evaluation. A larger question pool may improve task coverage, while repeated sampling can improve the stability of estimates on that pool; neither guarantees validity or personalization. The paper distinguishes the platform's current shared ranking from proposed task-specific and user-provided evaluations. Its central argument is that model selection requires evidence about performance on the intended work, not simply a high position on a general leaderboard.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑