发表机构
ETH Zurich; CSAIL, MIT; Helmholtz Zentrum München; University of Maryland, Baltimore County; Barcelona Supercomputing Center; Google; Duke Kunshan University; ETH AI Center(苏黎世联邦理工学院; 麻省理工学院计算机科学与人工智能实验室; 亥姆霍兹慕尼黑研究中心; 马里兰大学巴尔的摩县分校; 巴塞罗那超级计算中心; 谷歌; 昆山杜克大学; 苏黎世联邦理工学院人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出首个Croissant元数据提取基准(602篇论文),评估多种系统,发现单次提取优于智能体架构,尤其在长格式RAI字段差距最大,并发布基准、代码和排行榜。
AI 中文摘要
Croissant已成为机器学习可读数据集元数据的标准,然而填充其字段仍需要大量人工,并需仔细阅读随附的数据集文档。我们提出了首个基准测试,能够对符合社区标准模式的元数据提取进行端到端评估。该基准包含602篇论文,其中102篇带有经人工验证的金标准标注,500篇带有由大语言模型生成的银标准标注,覆盖完整的Croissant模式,包括核心字段和负责任人工智能(RAI)字段。利用该基准,我们评估了多种提取系统,涵盖前沿模型、开放权重模型和智能体架构,并采用双层评估框架,该框架结合了基于规则的评分和通过人工审计选出的LLM评判器。我们发现,单次提取始终优于我们评估的四种智能体架构:在各种骨干模型上,这些分解变体的准确率均低于单次全上下文传递。最大差距出现在长格式的RAI字段上,这些字段需要综合和解释分散在论文中的信息,而非从单一位置复制,而当前系统在此类场景下仍远非可靠。我们发布了该基准、评估代码、评判器审计、实时演示以及向新系统开放的排行榜。
英文摘要
Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation. We present the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema. The benchmark comprises 602 papers, including 102 with human-validated gold annotations and 500 with LLM-generated silver annotations, covering the full Croissant schema with both core and Responsible AI (RAI) fields. Using this benchmark, we evaluate a range of extraction systems spanning frontier models, open-weight models, and agentic architectures, under a two-tier evaluation framework that combines rule-based scoring with an LLM judge selected via human audit. We find that single-pass extraction consistently outperforms the four agentic architectures we evaluate: across backbones, these decomposed variants achieve lower accuracy than a single full-context pass. The largest gap appears on long-form RAI fields, which require synthesizing and interpreting information scattered across a paper rather than copying it from a single location, a setting where current systems remain far from reliable. We release the benchmark, evaluation code, judge audit, a live demo, and a leaderboard open to new systems.
CommentsAccepted at NeurIPS 2026 (Track on Evaluations and Datasets). Website: https://berkearda.github.io/croissantminer/