arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SetGo:科学人工智能数据集的元数据就绪性

SetGo: Metadata Readiness for Scientific AI Datasets

Sean R. Wilkinson, Polina Shpilker, Wesley Brewer

arXiv 2607.22677首次发表:更新:

发表机构

Oak Ridge National Laboratory(橡树岭国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对科学AI数据集元数据就绪性评估工具缺失的问题,提出开源工具SetGo,能从六个维度评估修复元数据,揭示通用工具未发现的缺陷,提升公平性分数,还可集成大语言模型支持自动化工作流程。

AI 中文摘要

用于人工智能的科学数据集需要计算就绪性以进行模型训练,以及元数据就绪性以实现发现、共享和重用。数据集成就绪引擎(REDI)解决了计算就绪性问题,但没有相应工具评估数据集的元数据是否足够完整、受治理且符合标准以用于发布和基于代理的使用。现有公平性评估工具仅对已发布的存储库记录进行操作,没有单个系统能同时涵盖公平性合规、许可、出处、治理、可重复性和目录就绪性。我们提出了SetGo,一个开源的Python工具包,可在数据集发布或存档之前评估和修复这六个维度的元数据就绪性。应用于四个科学语料库时,SetGo揭示了通用工具未检测到的缺陷。引导式丰富将整体公平性分数从52 - 57%提高到81 - 91%,并且单个setgo发布命令可将数据集推送到Hugging Face Hub、CKAN或带有ML Commons Croissant 1.0元数据边车的OpenMetadata。为了支持交互式和自动化工作流程,SetGo通过/setgo技能与由大语言模型(LLMs)驱动的编码代理集成,实现自然语言执行完整的评估 - 丰富 - 发布循环,用户只需提供缺失的元数据值。

英文摘要

Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, sharing, and reuse. The Readiness Engine for Data Integration (REDI) addresses computational readiness, but no corresponding tool evaluates whether a dataset's metadata are sufficiently complete, governed, and standards-compliant for publication and agent-based consumption. Existing FAIR assessors operate only on published repository records, and no single system covers FAIR compliance, licensing, provenance, governance, reproducibility, and catalog readiness together. We present SetGo, an open-source Python toolkit that assesses and repairs metadata readiness across these six dimensions before a dataset is published or archived. Applied to four scientific corpora, SetGo surfaces deficiencies that general-purpose tools do not detect: ERA5 climate metadata scores 4% on ACDD 1.3 compliance; materials datasets fail OPTIMADE species-definition requirements; and PDB-derived proteomics data carries licensing terms incompatible with standard SPDX identifiers. Guided enrichment raises overall FAIR scores from 52-57% to 81-91%, and a single setgo publish command pushes to Hugging Face Hub, CKAN, or OpenMetadata with ML Commons Croissant 1.0 metadata sidecars. To support interactive and automated workflows, SetGo integrates with coding agents powered by large language models (LLMs) through a /setgo skill that enables natural-language execution of the full assess-enrich-publish loop, with user involvement limited to supplying missing metadata values.

Comments6 pages; accepted at SSDBM 2026

DOI:10.1145/3828820.3828827

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑