AI 中文总结
本研究通过系统分析14,767篇论文,映射了大语言模型基准设计的变化趋势,发现评估日益强调行动、交互与专业应用,并质疑扩展评估是否提供独立证据或复制模型偏好。
AI 中文摘要
基准测试是评估和传达大语言模型(LLMs)进展的核心方式。然而,仅凭模型排名几乎无法揭示评估需求本身的变化。不断扩展的基准种类提供了另一个视角:研究人员期望大语言模型做什么,以及他们将什么视为成功表现。我们系统地梳理了2022年1月至2026年8月期间arXiv提交中引入或更新评估资源的14,767篇论文。通过分阶段筛选和自动化全文编码,我们考察了目标系统与领域、评估材料与条件以及评分机制的变化。该集合显示,对行动、交互和专业应用的重视日益增强,而既有和较新的设计元素也频繁共存。模型参与度的发展亦不均衡:基于LLM的评分在智能体组和非智能体组中均有所增长,而模型生成的材料在近期队列中未显示出类似的持续增长。这些发现揭示了公共研究如何将能力期望转化为具体的测试和成功标准。随着AI参与构建测试、执行任务和评判响应,它们也引发了一个问题:不断扩展的评估是否提供了更独立的证据,还是有可能复制其参与模型的偏好和盲点?
英文摘要
Benchmarks are central to how progress in large language models (LLMs) is assessed and communicated. Yet model rankings alone reveal little about how evaluation requirements themselves are changing. The expanding variety of benchmarks offers another perspective: what researchers expect LLMs to do, and what they count as successful performance. We systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026. Using staged screening and automated full-text coding, we examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms. The collection shows growing emphasis on action, interaction, and professional applications, while established and newer design elements frequently coexist. Model participation also develops unevenly: LLM-based scoring grows within both agent and non-agent groups, whereas model-generated materials show no comparable sustained increase in recent cohorts. These findings illuminate how public research translates capability expectations into concrete tests and criteria for success. As AI participates in constructing tests, performing tasks, and judging responses, they also raise a question: does expanding evaluation provide more independent evidence, or risk reproducing the preferences and blind spots of its participating models?
Comments15 pages, 5 figures, 7 tables. Data and code: https://github.com/xxcg322/LLM-Bench-Map