文献引导的基于描述符的组成搜索空间过滤框架
A literature-guided descriptor-based framework for filtering composition search spaces
浏览论文内容
中文总结 AI 辅助
提出文献引导的描述符过滤框架,利用词嵌入选择描述符构建帕累托过滤器,平均过滤74.27%候选组成,误差1.93%,平衡保留比例与性能。
中文摘要 AI 辅助
科学文献中包含关于材料行为的潜在知识,但其中大部分知识是通过词语、语境和反复出现的关联来表达的,而非明确的设计原则。这引出了一个核心问题:如何将大规模科学语料库用于材料发现中的实际问题?在此,我们提出了一种文献引导的基于描述符的过滤框架,用于缩减组成搜索空间。对于给定的性能指标,该框架从文献训练的词嵌入模型中的过滤词汇表中选取两个描述符,并利用所选描述符构建基于帕累托的过滤器,以筛选候选组成。在所评估的性能指标和组成搜索空间中,该框架平均过滤掉74.27%的候选组成,相对于实验测量值的平均最优值误差为1.93%。与专家选择和随机描述符相比,我们的依赖于性能指标的描述符在保留比例和最优值误差之间提供了更可控的平衡。这些结果表明,文献衍生的嵌入可以支持直观且可复现的过滤器,以缩小候选组成空间,同时保留高性能组成。
英文摘要
Scientific literature contains latent knowledge about materials behavior, but much of this knowledge is expressed through words, contexts, and recurring associations rather than explicit design principles. This raises a central question: how can large-scale scientific corpora be used for practical problems in materials discovery? Here, we present a literature-guided descriptor-based filtering framework for reducing composition search spaces. For a given performance metric, the framework selects two descriptors from a filtered vocabulary in a literature-trained word embedding model and uses the selected descriptors to construct a Pareto-based filter for candidate compositions. Across the evaluated performance metrics and composition search spaces, the framework filters out an average of 74.27\% of the candidate compositions, with an average best-value error of 1.93\% relative to experimental measurements. Compared with expert-chosen and random descriptors, our performance metric-dependent descriptors provide a more controlled balance between retained fraction and best-value error. These results show that literature-derived embeddings can support intuitive and reproducible filters for narrowing candidate composition spaces while preserving high-performing compositions.