AI 中文总结
该研究通过在6个虚拟分子库上对4种分子语言模型进行基准测试,发现显式领域适配可提升分子表示性能,适配后的编码器在基准任务中表现最佳,为虚拟筛选等领域提供了样本高效的自适应决策策略。
AI 中文摘要
预训练分子语言模型正日益被用作分子编码器,以学习结构-性能关系。然而,它们在其预训练领域内外用于分子发现的实际适用性仍不明确。本文中,我们系统地在涵盖药物发现、有机材料和催化的6个虚拟分子库上对4种分子语言模型进行基准测试。原生分子语言模型嵌入在不同库中的发现性能存在显著差异,而分子指纹提供了始终强劲且稳健的基线。与潜在的领域-表示不匹配一致,我们表明显式领域适配可显著提升表示性能。对目标虚拟库的结构进行微调分子语言模型编码器可始终提高样本效率,多个适配后的编码器成为基准任务中表现最佳的表示。这些结果表明,分子表示质量强烈依赖于目标领域,显式适配可提升分子基础模型的实际实用性。更广泛地说,我们的发现确立了领域适配的分子表示是虚拟筛选和自驱动实验室中样本高效自适应决策的有前景策略。
英文摘要
Pretrained molecular language models are increasingly used as molecular encoders for learning structure-property relationships. However, their practical suitability for molecular discovery within and beyond their pretraining domain remains unclear. Herein, we systematically benchmark four molecular language models across six virtual molecular libraries spanning drug discovery, organic materials, and catalysis. Native molecular language model embeddings show substantial variation in discovery performance across libraries, whereas molecular fingerprints provide a consistently strong and robust baseline. Consistent with a potential domain-representation mismatch, we show that explicit domain adaptation substantially improves representation performance. Fine-tuning molecular language model encoders on structures from the target virtual library consistently improves sample efficiency, with several adapted encoders emerging as the top-performing representations across the benchmark tasks. These results show that molecular representation quality depends strongly on the target domain and that explicit adaptation can improve the practical utility of molecular foundation models. More broadly, our findings establish domain-adapted molecular representations as a promising strategy for sample-efficient adaptive decision making in virtual screening and self-driving laboratories.