arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

科学数据技能:实现规模化的智能体就绪型科学数据服务

Scientific Data Skills: Enabling Agent-Ready Scientific Data Services at Scale

Xiaohan Huang, Qingqing Long, Xiaolei Du, Siyu Pu, Jiawen Xu, Haotian Chen, Chenyang Zhao, Jinbiao Liu, Xuezhi Wang, Hengshu Zhu, Yuanchun Zhou

arXiv 2608.19625首次发表:更新:

发表机构

Computer Network Information Center, Chinese Academy of Sciences; University of the Chinese Academy of Sciences(中国科学院计算机网络信息中心; 中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有数据集表示难以满足AI智能体需求的问题,提出SciDSK智能体就绪型表示并构建相关平台,可提升智能体的数据集发现与解释能力。

AI 中文摘要

科学数据正被AI智能体日益广泛地使用,但现有数据集表示对自主发现、解释和调用的支持有限。这一限制源于科学数据分散在异构存储库中,且现有数据集表示主要为人类使用而设计。为解决这一问题,我们提出科学数据技能(SciDSK),一种智能体就绪型表示,将数据集特定知识和操作指导打包为可复用的智能体技能。SciDSK整合了数据集描述、科学背景、文件组织、使用流程、质量检查和来源信息,同时将底层数据保留在原始存储库中。我们定义了结构化的SciDSK规范,并开发了系统化的构建流程,该流程将每个SciDSK建立在权威数据集记录及相关支撑材料的基础上。我们还建立了科学数据技能库,这是一个统一平台,可发布涵盖六个科学领域的SciDSK资源,并支持包访问、持久标识和对源数据集的可追溯性。我们通过数据集发现的检索基准和数据集解释的受控案例对SciDSK进行评估。结果表明,SciDSK提升了智能体驱动的数据集发现能力,并为数据集解释提供了更精确、可操作的支持。这些发现证实了以智能体就绪型表示组织数据集特定知识的价值。

英文摘要

Scientific data are increasingly used by AI agents, yet existing dataset representations provide limited support for reliable dataset discovery and interpretation, constraining their effective use in scientific workflows. This limitation arises because agents must search across heterogeneous repositories and reconstruct dataset-specific semantics and operating procedures from documentation designed primarily for human use. To address this limitation, we introduce the Scientific Data Skill (SciDSK), an agent-ready representation that packages dataset-specific knowledge and operational guidance as a reusable agent skill. A SciDSK integrates dataset descriptions, scientific context, file organization, task-specific usage procedures, quality checks, and provenance information while retaining the underlying data in its original repository. We define a structured SciDSK specification and develop a systematic construction pipeline that grounds each SciDSK in authoritative dataset records and associated supporting materials. We further establish the Scientific Data Skill Bank, a unified platform that publishes SciDSK resources across six scientific disciplines and supports package access, persistent identification, and traceability to source datasets. We evaluate SciDSK through a retrieval benchmark for dataset discovery and controlled cases for dataset interpretation. On the query retrieval benchmark, Agent-SciDSK achieves 80.77% Hit@1, exceeding Agent-Raw by 9.62 percentage points. Across controlled interpretation cases, the SciDSK condition satisfies 23 of 24 assessment criteria, compared with 22 under the web-page condition. These results indicate that SciDSK improves how agents locate and understand scientific datasets, providing a stronger foundation for actionable scientific data use.

Comments15 pages, 5 figures, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑