arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ERAF4XRD:一种从科学文献构建经过验证的实验X射线衍射数据库的多模态智能体框架

ERAF4XRD: A multimodal agentic framework for constructing validated experimental X-ray diffraction databases from scientific literature

Afnan Mostafa, William Ratcliff, Simon J. L. Billinge, Niaz Abdolrahim

arXiv 2609.18583首次发表:更新:

发表机构

University of Rochester; National Institute of Standards and Technology; University of Maryland, College Park; University of California, Santa Barbara; Laboratory for Laser Energetics, University of Rochester(罗切斯特大学; 美国国家标准与技术研究院; 马里兰大学帕克分校; 加州大学圣塔芭芭拉分校; 罗切斯特大学激光能学实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ERAF4XRD是一个全自动多模态多智能体框架,从科学文献中识别XRD图形、提取并验证元数据,重建经过验证的实验XRD数据库,在273篇论文基准上达到98.7%识别准确率和98.5%验证精确率。

AI 中文摘要

科学文献中包含数十年的实验测量数据,但这些数据作为结构化数据难以被现代AI和数据驱动研究所获取。这些信息大部分分布在图表、标题、正文和表格中,需要先识别、关联并验证实验数据及其背景信息,才能加以重用。在此,我们介绍了ERAF4XRD(实验读取器智能体框架,用于X射线衍射),这是一个全自动的多模态(即图像和文本)多智能体框架,可从科学出版物中重建经过验证的X射线衍射(XRD)记录。ERAF4XRD下载并筛选文档,识别XRD图形,提取元数据并将其链接到相应的实验数据,并使用独立的验证智能体根据源证据验证输出。在一个包含273篇科学出版物和3150个候选图形的人工整理基准上,ERAF4XRD在XRD图形识别方面达到了高达98.7%的准确率,并生成了跨22个字段的1400个元数据值。对最终验证记录的独立人工评估显示,精确率为98.5%,召回率为90.7%,在443个评估字段中未观察到无支撑的元数据。通过超越信息提取,转向关联实验记录的重建和验证,ERAF4XRD建立了一种自动化方法,将已发表的科学信息转化为可供AI和数据驱动科学使用的机器可读实验数据集。

英文摘要

The scientific literature contains decades of experimental measurements that remain difficult to access as structured data for modern AI and data-driven research. Much of this information is distributed across figures, captions, text, and tables, requiring experimental data and their context to be identified, connected, and verified before they can be reused. Here we introduce ERAF4XRD (Experiment Reader Agentic Framework for X-Ray Diffraction), a fully automated multimodal (i.e., image and text), multi-agent framework that reconstructs validated X-ray diffraction (XRD) records from scientific publications. ERAF4XRD downloads and screens documents, identifies XRD figures, extracts and links metadata to the corresponding experimental data, and validates outputs against source evidence using an independent validation agent. On a manually curated benchmark of 273 scientific publications containing 3,150 candidate figures, ERAF4XRD achieved up to 98.7% accuracy for XRD figure identification and generated 1,400 metadata values across 22 fields. Independent manual assessment of the final validated records yielded 98.5% precision and 90.7% recall, with no unsupported metadata observed among 443 evaluated fields. By moving beyond information extraction to the reconstruction and validation of linked experimental records, ERAF4XRD establishes an automated approach for transforming published scientific information into machine-readable experimental datasets for AI and data-driven science.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑