晶体生成器离散组合信道上的强化学习:验证增益与奖励黑客
Reinforcement Learning on the Discrete Composition Channel of a Crystal Generator: Validated Gains and Reward Hacking
浏览论文内容
中文总结 AI 辅助
本研究用GRPO强化学习对齐晶体生成模型,通过直接优化原子类型转移似然,将mSUN产出率从13.4%提升至45.5%,并发现奖励黑客问题,提出按参考相数量拆分报告的基准改进建议。
中文摘要 AI 辅助
逆向材料设计是计算材料发现领域长期以来的目标。晶体材料的生成模型通常被训练为匹配结构数据库的分布,而其训练目标中没有任何内容指向特定的设计目标,如目标性能。我们使用组相对策略优化(GRPO)将基于随机插值和离散流匹配的生成模型与通用黑盒奖励函数通过强化学习对齐。原子类型由离散流生成,我们推广的GRPO的策略梯度直接作用于原子类型转移的似然,这使我们的工作区别于先前针对晶体材料扩散和基于流的生成模型的强化学习方法。我们引入一个奖励函数,将亚稳态、独特且新颖的结构(mSUN)的产出率从预训练模型的13.4%提高到强化模型的45.5%,这是由社区基准评估的。我们的奖励还提高了基于潜在去噪扩散模型的晶体材料强化学习框架的性能。同时,我们发现直接强化原子类型转移似然会导致奖励利用,必须通过显式防护来阻止。同样的分析也暴露了社区指标中的差距。不同堆积中的单元素结构被计为亚稳态、独特且新颖的材料,从而虚增mSUN,而不会产生任何新化合物。稳定性声明的好坏取决于其参考包络。我们按每个结果背后的参考相数量来拆分报告所有结果,并主张基准也应如此。
英文摘要
Inverse materials design is a long-standing goal of computational materials discovery. Generative models for crystalline materials are typically trained to match the distribution of a structure database, while nothing in their training objective points them at specific design goals such as targeted properties. We use group-relative policy optimization (GRPO) to align a generative model based on stochastic interpolants and discrete flow matching with general black-box reward functions through reinforcement learning. Atom types are generated by a discrete flow and the policy gradient of our generalization of GRPO directly acts on the likelihoods of the atom-type transitions, which differentiates our work from previous reinforcement-learning approaches for diffusion and flow-based generative models of crystalline materials. We introduce a reward function that raises the yield of metastable, unique and novel structures (mSUN) from 13.4% for the pretrained model to 45.5% for the reinforced model, as evaluated by a community benchmark. Our reward also improves the performance of a reinforcement learning framework for crystalline materials based on latent denoising diffusion models. At the same time, we find that directly reinforcing atom-type transition likelihoods enables reward exploitation that has to be prevented with explicit guards. The same analysis also exposes a gap in the community metric. Single-element structures in distinct packings are counted as metastable, unique and novel materials and inflate mSUN without yielding any new compounds. A stability claim is only as good as its reference hull. We report every result split by the number of reference phases behind it and argue that benchmarks should do the same.
发表机构
- University of Florida(佛罗里达大学)
- New York University(纽约大学)
- Oak Ridge National Laboratory(橡树岭国家实验室)
机构由 AI 辅助整理,请以论文原文为准。