让CSP成为你的ANCHOR:基于冻结结构先验的自适应晶体搜索
Let CSP Be Your ANCHOR: Adaptive Crystal Search over Frozen Structure Priors
浏览论文内容
中文总结 AI 辅助
ANCHOR通过分离组成搜索与结构生成,利用GRPO策略和连续自适应新颖性,在冻结CSP先验上显著提升晶体发现的新颖性与稳定性指标。
中文摘要 AI 辅助
从头晶体生成(DNG)模型决定在组成空间中搜索何处以及如何用一组权重生成结构。我们认为,将这两者分离能更好地服务于发现。晶体结构预测(CSP)模型是一个物理先验,应通过似然训练来改进,而奖励(包括针对搜索自身历史测量的新颖性)应作用于对组成的搜索。我们引入了ANCHOR,一种围绕冻结CSP模型、以多目标奖励训练的GRPO组成策略,以及连续自适应新颖性(CAN),一种针对已知结构和不断增长的发现历史的分级新颖性评分。在单一评估器下使用冻结的CSP模型作为固定标尺,我们测试了自适应应作用于何处。在相同的CSP骨干网络上,用ANCHOR的策略替换DNG组成,在99.9%的化学式唯一性下,MSUN从11.4%提高到47.6%,SUN从1.1%提高到22.1%。直接在相同奖励上微调DNG模型,只会移动其组成边际,而不会提高其在壳上的比例。我们证明,对DNG模型进行KL正则化微调只能以有界因子重新加权预训练模型已支持的化学组成,而无正则化的DNG微调则趋向于已知或较不稳定的化学组成。即使仅将稳定性奖励路由到ANCHOR的CSP骨干网络中,相对于冻结骨干网络,SUN也大致减半,而在搜索过程中发现的结构上进行似然训练可以改进CSP骨干网络。在MatterGen的评估流程下,ANCHOR将最先进的MSUN从29.2%提高到41.3%,无需重新训练即可迁移到另外两个CSP骨干网络,并在蒸馏到Crystalite-CSP后达到47.1%。与任何针对势优化的模型一样,其在壳上的比率取决于该势。
英文摘要
De novo crystal generation (DNG) models decide where to search in composition space and how to generate structures with one set of weights. We argue that discovery is better served by separating the two. A crystal structure prediction (CSP) model is a physical prior that should be improved by likelihood training, while rewards, including novelty measured against the search's own history, should act on a search over compositions. We introduce ANCHOR, a GRPO composition policy trained with multi-objective rewards around a frozen CSP model, and continuous adaptive novelty (CAN), a graded novelty score against known structures and a growing discovery history. Using the frozen CSP model as a fixed ruler under one evaluator, we test where adaptation should act. Replacing DNG compositions with ANCHOR's policy on the same CSP backbone raises MSUN from 11.4% to 47.6% and SUN from 1.1% to 22.1% at 99.9% formula uniqueness. Fine-tuning DNG models directly on the same rewards instead moves their composition marginal without raising their on-hull fraction. We show that KL-regularized fine-tuning of a DNG model can only reweight chemistry the pretrained model already supports by a bounded factor, while unregularized DNG fine-tunes move toward known or less stable chemistry. Even a stability-only reward routed into ANCHOR's CSP backbone roughly halves SUN relative to the frozen backbone, whereas likelihood training on structures found during search can improve a CSP backbone. Under MatterGen's evaluation pipeline, ANCHOR raises state-of-the-art MSUN from 29.2% to 41.3%, transfers without retraining to two further CSP backbones, and reaches 47.1% after distillation into Crystalite-CSP. As with any model optimised against a potential, its on-hull rate depends on that potential.
发表机构
- Technical University of Denmark(丹麦技术大学)
- University of Southern Denmark(南丹麦大学)
机构由 AI 辅助整理,请以论文原文为准。