arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基准雷达:面向AI基准与评估的活数据库和搜索引擎

Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation

Koutian Wu, Junjie Zhou, Ergan Shang, Jiayu Wang, Pengqian Han, Junkai Wang, Wanghan Xu, Songyuanyi Lu, Lin Shi

arXiv 2609.11115首次发表:更新:

发表机构

Earth-Space-AI; Tacite AI; Hangzhou Dianzi University; Carnegie Mellon University; The University of Auckland; Tsinghua University; Shanghai Jiao Tong University(地球空间人工智能; Tacite AI; 杭州电子科技大学; 卡内基梅隆大学; 奥克兰大学; 清华大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

基准雷达是一个活数据库和搜索引擎,通过每日发现和目录检索,帮助研究人员查找和比较AI基准,并提供排行榜和帕累托前沿等分析工具。

AI 中文摘要

基准研究人员以及大型语言模型(LLM)和其他AI系统的开发者需要找到相关的评估,定位其基准数据集和代码,并理解报告分数背后的设置。我们提出了基准雷达(Benchmark Radar),这是一个用于检索和发现AI基准的活数据库和搜索引擎,涵盖LLM评估、智能体(agentic)和工具使用基准、编程、推理、安全以及领域特定评估。该系统将基准论文、代码库、数据集和发布的每日发现与可搜索的基准目录、模型卡和技术报告中的提及以及分数历史相结合。它保留来源身份和引用,以便读者可以检查候选基准及其评估证据。每日发现依赖37个来源:13个直接连接器和24个第一方研究与工程信息流。目录包含来自4个基准目录的1,283条来源记录,以及790条记录上的12,916个数值观测。我们描述了收集和检索过程,审计了完整目录,并考察了基准饱和、采用趋势以及分数比较的局限性。一个工作示例完整演示了现有技术搜索,展示了在设计新评估时如何查询目录和检查基准证据。我们发布了带有基准排行榜、分数与实际使用情况的帕累托前沿视图、饱和度和趋势视图、每日信息流、可下载证据、用于离线查询的命令行界面(CLI)以及可复现分析的Web仪表板。

英文摘要

Benchmark researchers and developers of large language models (LLMs) and other AI systems need to find relevant evaluations, locate their benchmark datasets and code, and understand the settings behind reported scores. We present Benchmark Radar, a living database and search engine for retrieval and discovery of AI benchmarks, covering LLM evaluation, agentic and tool-use benchmarks, coding, reasoning, safety, and domain-specific evaluations. The system combines daily discovery of benchmark papers, repositories, datasets, and releases with a searchable benchmark catalog, mentions in model cards and technical reports, and score histories. It retains source identities and citations so readers can inspect candidate benchmarks and their evaluation evidence. Daily discovery draws on 37 sources: 13 direct connectors and 24 first-party research and engineering feeds. The catalog contains 1,283 source records drawn from 4 benchmark catalogs and 12,916 numeric observations on 790 records. We describe collection and retrieval, audit the full catalog, and examine benchmark saturation, adoption trends, and the limits of score comparisons. A worked example walks through a complete prior-art search, showing how to query the catalog and inspect benchmark evidence when designing a new evaluation. We release the web dashboard with a benchmark leaderboard, a Pareto frontier view of score against measured use, saturation and trend views, daily feeds, downloadable evidence, a command-line interface (CLI) for offline queries, and reproducible analysis.

CommentsCode: https://github.com/ktwu01/benchmark-radar, Project site: https://benchmark-radar.org

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑