通过电池文献的多模态挖掘利用X射线吸收光谱数据
Harnessing X-ray Absorption Spectroscopy Data through Multimodal Mining of Battery Literature
- Argonne National Laboratory(阿贡国家实验室)
- NVIDIA(英伟达)
- University of California, Berkeley(加州大学伯克利分校)
- Lawrence Berkeley National Laboratory(劳伦斯伯克利国家实验室)
- The University of Chicago(芝加哥大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究利用多模态文献挖掘将电池文献中分散的X射线吸收光谱数据转化为可用于人工智能的实验数据资源,开发数字化管道并应用于电池文献,产生开放数据集,为相关分析、比较及材料发现提供基础。
AI中文摘要:
X射线吸收光谱(XAS)对于理解材料的局部电子和原子结构至关重要,但大多数已发表的光谱因嵌入在文献中的图表及碎片化文本描述而难以进行数据驱动分析。本文利用多模态(图像和文本)文献挖掘将这些分散的知识转化为可用于人工智能的实验数据资源。开发了可扩展的光谱数据数字化管道,识别全文文章中的XAS图表,数字化光谱曲线,并将每个光谱与测量边缘和材料的相关元数据链接。应用该管道于电池文献产生了包含13740个XAS光谱的开放数据集,涵盖66个吸收元素和多种电池化学组成,经专家验证确认了光谱和元数据信息的准确提取。通过将文献中嵌入的光谱转化为结构化数值数据,该数据集为大规模XAS分析、跨实验室比较、高通量表征及先进材料的自主发现提供了基础。
英文摘要:
X-ray absorption spectroscopy (XAS) is central to understanding the local electronic and atomic structure of materials, yet most published spectra remain inaccessible to data-driven analysis because they are embedded in figures and described through fragmented textual context in the literature. Here, we use multimodal (image and text) literature mining to transform this dispersed knowledge into an AI-ready experimental data resource. We developed a scalable spectroscopy data digitization pipeline that identifies XAS figures in full-text articles, digitizes spectral curves, and links each spectrum to accompanying metadata on the measured edge and material. Applying this pipeline to the battery literature produced an open dataset of 13,740 XAS spectra, spanning 66 absorbing elements and diverse battery chemistries, with expert validation confirming accurate extraction of spectral and metadata information. By converting literature-embedded spectra into structured numerical data, this dataset provides a foundation for large-scale XAS analysis, cross-laboratory comparison, high-throughput characterization, and autonomous discovery of advanced materials.