发表机构
Centre for Credible AI; Warsaw University of Technology; University College Cork; University of Technology Sydney; Human-Centered AI Lab; University of Pisa; ISTI-CNR; Technical University of Berlin; Fraunhofer Heinrich Hertz Institute; Berlin Institute for the Foundations of Learning and Data; University of Warsaw(可信人工智能中心; 华沙理工大学; 科克大学学院; 悉尼科技大学; 以人为中心的人工智能实验室; 比萨大学; 意大利国家研究委员会信息科学与技术研究所; 柏林工业大学; 弗劳恩霍夫海因里希·赫兹研究所; 柏林学习与数据基础研究所; 华沙大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出Modelpedia框架,用于从论文提取模型发现并聚合为可搜索目录,已在ICLR论文中提取超千条发现,助力AI元科学发展。
AI 中文摘要
关于AI模型的科学知识产生速度远超学界整理的速度。每隔几个月就会有一个新的基础模型重塑该领域,数百篇论文、博客和技术报告记录了每个模型的表现或缺陷,但这些发现仍分散且无法有效检索。为解决这一缺口,我们提出Modelpedia,一个自动化、大语言模型(LLM)辅助的框架,它从已发表论文中提取关于模型的发现,将其与涉及的模型、数据集、方法和概念关联,并将结果聚合为可搜索的公开目录。将原型应用于2024和2025年ICLR接收的论文,我们提取了超过1000条发现,并将该目录本身作为研究对象,对学界研究模型的方式进行元分析。现在,我们邀请学界探索、贡献并基于该开放目录开展工作,助力将模型发现确立为AI元科学的共同基础。
英文摘要
Scientific knowledge about AI models is produced faster than the community can organize it. Every few months a new foundation model reshapes the field and hundreds of papers, blogs, and technical reports document how each behaves or fails. Yet, these findings remain scattered and effectively unretrievable. To address this gap we present Modelpedia, an automated, LLM-assisted framework that extracts findings about models from published papers, links it to the model, dataset, method, and concept it concerns, and aggregates the result into a searchable public catalog. Applying the prototype to accepted ICLR 2024 and 2025 papers, we extract over a thousand findings and, treating the catalog itself as an object of study, run a meta-analysis of how the community investigates models. Now, we invite the community to explore, contribute to, and build on the open catalog, and to help establish model findings as a shared foundation for the meta-science of AI.
CommentsFor the website, see: https://credibleai.github.io/modelpedia For the codebase, see: https://github.com/CredibleAI/modelpedia