用于能源领域大语言模型应用的多模态数据集
A Multimodal Dataset for Large Language Model Applications in the Energy Domain
- UBITECH
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
介绍mAIEnergy多模态数据集,整合文本、图像、时间序列及地理空间等数据,统一为结构化格式并伴有元数据与工作流程,可作能源知识库,遵循FAIR原则,助力能源领域大语言模型应用的研究、建模和决策。
AI中文摘要:
本文介绍了mAIEnergy数据集,这是一个开放获取的多模态语料库,旨在支持能源领域的大语言模型应用。该数据集整合了约50000篇文本文件、20000张图像、2500万个数值时间序列记录以及200万个地理空间和关系数据条目。数据涵盖政策法规文本、科学文章、新闻文章、卫星和情境图像、电力系统测量、气象观测等。所有数据已统一为结构化、即用格式,并伴有一致的元数据及可重复的数据检索和准备工作流程。该数据集可作为基础能源知识库,遵循FAIR原则,增强了其在人工智能驱动的能源研究、建模和决策中的适用性。
英文摘要:
This paper presents the mAIEnergy dataset, an open-access, multimodal corpus developed to support Large Language Model (LLM) applications in the energy sector. The dataset integrates approximately 50,000 textual documents, 20,000 images, 25 million numerical time series records, and 2 million geospatial and relational data entries. It includes policy and regulatory texts, scientific articles and news articles, satellite and contextual imagery, electricity system measurements, weather observations, statistical indicators, and geospatial representations of energy infrastructure and related entities. All data have been harmonized into structured, ready-to-use formats, accompanied by consistent metadata and reproducible data retrieval and preparation workflows. The dataset can serve as a foundational energy knowledge base, allowing energy stakeholders to integrate additional open-source or proprietary data. The mAIEnergy dataset adheres to Findable, Accessible, Interoperable, and Reusable (FAIR) principles, enhancing its applicability for AI-driven energy research, modeling, and decision-making.