AI 中文总结
ChemReporter是一款模块化框架,可将异构大规模化学数据集转为统一可查询格式,经处理、查询、导出三个阶段生成MLIP可用训练数据,支持处理超内存规模数据集,已在GitHub和PyPI发布。
AI 中文摘要
训练集的质量和多样性是决定机器学习原子间势(MLIP)可靠性的关键因素,然而直接使用完整的大规模数据集往往既不切实际又存在冗余,因此智能数据选择至关重要。但目前存在一个主要瓶颈:缺乏用于统一访问、整理和对异构大规模化学数据集进行子采样的基础设施,这些数据集在结构、元数据和文件格式上存在巨大差异。我们通过ChemReporter解决了这一缺口,它是一个模块化、与方法无关的框架,可将任意分子和材料数据集转换为统一的、可查询的表示形式,并将结果直接导出为可用于MLIP训练的训练数据。ChemReporter在三个解耦的阶段运行:处理阶段,将原始数据集解析为分区的Apache Parquet存储库,该存储库丰富了结构、物理和化学元数据;查询阶段,通过CLI或Python API使用任意选择条件(从简单的物理约束到用户自定义的策略)过滤和采样该存储库;导出阶段,将选定的子集流式传输到HDF5文件中,可直接用于现代MLIP训练框架。在整个过程中,每个导出的数据点都可追溯到其原始源条目,并且在相同配置和查询数据库版本下,数据集导出可被可靠复现。由于数据以可查询的、基于磁盘的格式存储,ChemReporter可以处理远大于可用内存的数据集,使其能够在标准计算基础设施上扩展到包含数十亿个结构的数据集。ChemReporter在GitHub和PyPI上以Apache许可证2.0发布。
英文摘要
Training set quality and diversity are key determinants of the reliability of machine learning interatomic potentials (MLIPs), yet using massive datasets in full is often impractical and redundant, making intelligent data selection essential. A major bottleneck, however, is the lack of infrastructure for uniformly accessing, curating, and subsampling heterogeneous large-scale chemical datasets, which differ widely in structure, metadata, and file format. We address this gap with ChemReporter, a modular, method-agnostic framework that converts arbitrary molecular and materials datasets into a unified, queryable representation and exports the results directly into MLIP-ready training data. ChemReporter operates in three decoupled stages: processing, which parses raw datasets into a partitioned Apache Parquet repository enriched with structural, physical, and chemical metadata; querying, which filters and samples this repository via a CLI or Python API using arbitrary selection criteria, from simple physical constraints to custom, user-defined strategies; and exporting, which streams the selected subset into an HDF5 file ready for direct use in modern MLIP training frameworks. Throughout this process, every exported data point remains traceable to its original source entry, and dataset exports can be reliably reproduced given the same configuration and query database version. Because data is stored in a queryable, disk-backed format, ChemReporter can process datasets far larger than available memory, allowing it to scale to billion-structure datasets on standard compute infrastructure. ChemReporter is available on GitHub and PyPI under the Apache License 2.0.
Comments19 pages, 10 figures