arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向FAIR与AI就绪数据:杰斐逊实验室GlueX实验数据生态系统评估

Toward FAIR and AI-Ready Data: An Assessment of the GlueX Experiment's Data Ecosystem at Jefferson Lab

Anil Panta, Brad Sawatzky, Casey Morean, Dmitry Romanov, Douglas Higinbotham

arXiv 2609.34987首次发表:更新:

发表机构

Thomas Jefferson National Accelerator Facility(托马斯·杰斐逊国家加速器设施)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究以杰斐逊实验室GlueX实验为试点,提出数据需先满足FAIR原则方可AI就绪,并构建了基于确定性指标的AI就绪评估框架,通过评估数据生态系统与现有验证软件,给出配套建议。

AI 中文摘要

美国能源部的项目日益要求实验数据不仅被保存,而且能被人工智能(AI)系统使用,但目前尚无公认的标准来定义数据集何时达到AI就绪。我们认为,数据必须先满足FAIR(可查找、可访问、可互操作、可重用)原则,然后才能达到AI就绪,且AI就绪是一个需要独立评估的独特属性。以杰斐逊实验室D厅的GlueX实验为试点,我们绘制了数据生态系统图,依据FAIR原则对其进行评估,并定义了一个AI就绪框架,其标准由确定性指标支撑。我们还评估了现有的验证软件,以区分哪些可以用现成工具评估,哪些需要领域特定的开发。评估产生了一系列发现,每项发现都配有一条建议,其中一些已投入生产或已原型化,另一些则计划作为后续步骤。

英文摘要

Department of Energy programs increasingly require that experimental data be not only preserved but usable by artificial intelligence (AI) systems, yet no accepted standard defines when a dataset is AI-ready. We argue that data must be FAIR (Findable, Accessible, Interoperable, Reusable) before they can be AI-ready, and that AI readiness is a distinct property requiring its own assessment. Using the GlueX experiment in Hall D at Jefferson Lab as a pilot, we map the data ecosystem, assess it against FAIR, and define an AI-readiness framework whose criteria are backed by deterministic metrics. We also evaluate existing validation software to separate what can be assessed with off-the-shelf tools from what requires domain-specific development. The assessment yields a set of findings, each paired with a recommendation, some of which are already in production or prototyped and others planned as next steps.

Comments18 pages, 1 figure, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑