arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Citrine信息学:化学与材料开发平台

Citrine Informatics: Chemical & Materials Development Platform

Maxwell C. Venetos, Steven J. Brown, Kenneth Kroenlein, Steven K. Kauwe, James E. Saal, Marco Musto, Matthew D. Gerboth, Kyle D. Miller, Gregory J. Mulholland

arXiv 2607.25039首次发表:更新:

AI 中文总结

该研究针对数据驱动材料发现面临的实验数据难复用、模型性能评估不准及设计空间受限等问题,提出Citrine平台,经四个协作阶段及相关方法,在多个案例研究中实现实验工作量大幅减少,促进各层共同发展。

AI 中文摘要

数据驱动的材料发现有望缩短从发明到应用长达数十年的历史进程,但将个别成功转化为持续的工业发现计划仍很困难。存在三个障碍:实验数据稀缺、成本高且格式难以复用;传统精度指标在发现所需的外推条件下夸大模型性能;现实设计空间受物理、可制造性、供应和成本限制。我们展示了Citrine平台,它历经十多年开发,作为对这些障碍的综合应对措施,并在封闭的顺序学习循环中组织为四个协作阶段。第一阶段通过材料数据图形表达(GEMD)模型摄取数据并提取特征,该模型将过程历史、测量不确定性和来源视为一级特征。第二阶段构建具有良好校准不确定性的机器学习模型,包括相关目标的多变量预测区间,并用外推交叉验证和动态发现指标而非随机留出分割来验证。第三阶段将成分、物理、加工和经济约束直接编码到设计空间中,第四阶段应用具有不确定性感知采集函数的FUELS顺序学习框架在严格评估预算下导航大型约束空间。已发表的案例研究表明,相对于随机搜索,实验工作量减少了两到九倍,说明了数据、建模和设计空间层不断共同发展的堆栈。

英文摘要

Today the Citrine Platform regularly powers data-driven materials discovery across industries, having moved beyond one-off demonstrations into routine industrial practice. Getting there required solving a core set of recurring obstacles: experimental data are scarce, costly, and published in formats that resist reuse; conventional accuracy metrics overstate model performance under the extrapolative conditions that define discovery; and realistic design spaces are bounded by physics, manufacturability, supply, and cost. Developed over more than a decade as an integrated response to these obstacles, the Citrine Platform is organized as four cooperating stages within a closed sequential learning loop. Stage 1 ingests and featurizes data through the Graphical Expression of Materials Data (GEMD) model, which treats process history, measurement uncertainty, and provenance as first-class features. Stage 2 builds machine learning models with well-calibrated uncertainty, including multivariate prediction intervals for correlated objectives, and validates them with extrapolative cross-validation and dynamic discovery metrics rather than random held-out splits. Stage 3 encodes compositional, physical, processing, and economic constraints directly into the design space, and Stage 4 applies the FUELS sequential learning framework with uncertainty-aware acquisition functions to navigate large constrained spaces under tight evaluation budgets. Published case studies spanning organic semiconductors, autonomous nanoparticle synthesis, and benchmark optimization tasks demonstrate two- to nine-fold reductions in experimental effort relative to random search, illustrating a stack in which data, modeling, and design-space layers continuously co-evolve.

Comments61 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑