发表机构
Tsinghua University; Beijing Institute of Mathematical Sciences and Applications; Wuhan University; Shenzhen MSU-BIT University; MathonAI Team(清华大学; 北京应用数学科学研究院; 武汉大学; 深圳莫斯科国立大学与北京理工大学联合大学; MathonAI团队)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对OLED分子设计面临的挑战,提出基于因果语言模型的逆分子设计框架,采用多阶段策略,经基础模型建立、属性预测器微调、强化学习等,最终能有效探索OLED化学空间,生成新候选物。
AI 中文摘要
有机发光二极管(OLED)材料的开发面临着巨大化学空间、严格量子化学约束和标记数据稀缺等多重挑战。尽管OLED生成问题很重要,但针对该特定领域有效训练的模型很少。我们提出了一种基于因果语言模型的逆分子设计框架,给定目标光电特性,模型直接生成满足特定约束的OLED SMILES序列。采用多阶段策略,先建立基础化学语言模型,再微调属性预测器,接着进行强化学习,最后通过DFT验证,表明该框架能有效探索OLED化学空间,生成具有高结构有效性和优化光电特性的新候选物。
英文摘要
The development of organic light-emitting diode (OLED) materials faces the compounded challenges of an astronomically large chemical space, stringent quantum-chemical constraints, and a scarcity of labeled data. Although the question of OLED generation is important, few models have been trained effectively for this specific domain. We propose an inverse molecular design framework based on causal language models: given target optoelectronic properties (e.g., excitation energy, oscillator strength), our model directly generates OLED SMILES sequences satisfying the specified constraints. We employ a multi-stage strategy: first, we establish a foundational chemical language model using a LLaMA-style transformer architecture. To the best of our knowledge, this represents the first successful adaptation of LLMs specifically for the OLED domain, bridging the gap between generic molecular generation and the stringent structural requirements of optoelectronic materials. Second, we fine-tune property predictors based on a BERT model pre-trained on our large-scale OLED dataset. Then, we perform Reinforcement Learning on our fine-tuned model, leveraging our property predictor, for better SMILES generation. Finally, through DFT verification, we demonstrate that our framework can efficiently navigate the OLED chemical space, generating novel candidates with high structural validity and optimized optoelectronic properties.