将化学知识编译为可执行描述符用于材料预测
Compiling Chemical Knowledge into Executable Descriptors for Materials Prediction
浏览论文内容
中文总结 AI 辅助
本研究提出CRISP框架,将化学知识编译为可执行描述符用于材料预测,其在无机晶体可合成性等任务上表现优于现有方法,建立了数据集无关的可迁移表征途径。
中文摘要 AI 辅助
材料预测高度依赖于科学知识的表征方式,然而许多关键的调控因素仅以自然语言启发式的形式存在,无法被传统学习器利用。我们提出CRISP,一种大语言模型辅助框架,该框架将表征构建视为规则空间探索与编译问题:在不获取结构、标签或数据划分的情况下,反复采样与目标相关的化学规则,整合相关概念并将其编译为可执行的标量描述符,供传统学习器使用。针对正-未标记无机晶体可合成性任务,在相同学习器下,CRISP的表现优于专家定制及通用结构表征,且超过专用可合成性模型;在结构规模与化学家族发生变化时,其优势最为显著。生成频率较低的规则贡献了互补的预测信息,表明生成频率不决定其效用。相同工作流为形成能与离子电导率生成了具有竞争力的表征,同时揭示了剪切模量的任务相关限制,建立了从广泛化学知识到可迁移计算表征的数据集无关、可审计的途径。
英文摘要
Materials prediction depends critically on how scientific knowledge is represented, yet many governing considerations exist only as natural-language heuristics that conventional learners cannot use. We introduce CRISP, a large language model-assisted framework that treats representation construction as a rule-space exploration and compilation problem: it repeatedly samples target-relevant chemical rules without access to structures, labels or data splits, consolidates related concepts, and compiles each into an executable scalar descriptor supplied to a conventional learner. For positive-unlabeled inorganic-crystal synthesizability, CRISP outperformed expert-curated and generic structural representations under a shared learner and surpassed purpose-built synthesizability models, with its advantage most pronounced under structural-size and chemical-family shifts. Infrequently generated rules contributed complementary predictive information, showing that generation frequency does not determine utility. The same workflow yielded competitive representations for formation energy and ionic conductivity while revealing task-dependent limits for shear modulus, establishing a dataset-blind, auditable route from broad chemical knowledge to transferable computational representations.