语言模型能否设计出结合分子?在空间约束下对语言模型进行基准测试
Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints
浏览论文内容
中文总结 AI 辅助
研究探讨通用语言模型在3D分子设计中应对复杂空间约束的能力,引入3D - Fit评估策略,发现语言模型虽落后于先进方法,但有潜力同时处理多种空间约束。
中文摘要 AI 辅助
基于结构的药物设计(SBDD)利用蛋白质靶点的3D结构(常辅以其他空间约束)来生成候选结合分子。虽然扩散模型是高质量3D分子生成的主导范式,但基于语言模型(LLM)的方法在分子设计中迅速兴起,在口袋条件分子生成中表现出竞争力。然而,其处理物理和3D空间环境的能力尚未充分探索。本文系统分析了当前通用LLM与专业扩散模型等基线相比,是否能应对复杂3D约束。考虑了基于蛋白质口袋的3D配体生成及相关空间约束,引入3D - Fit评估策略。结果表明LLM虽仍落后于先进方法,但有潜力且能同时处理多种空间约束。
英文摘要
Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to generate candidate binding molecules. While diffusion models have dominated as a leading paradigm for high-quality 3D molecule generation, LLM-based methods are rapidly emerging in molecular design and have shown competitive performance in pocket-conditioned molecular generation. However, their ability to reason about physics and 3D spatial environments is largely underexplored. In this work, we systematically analyze whether current general-purpose LLMs are capable of navigating complex 3D constraints compared to established baselines such as specialized diffusion models. We consider 3D ligand generation conditioned on protein pockets together with ligand- and interaction-derived spatial constraints, including anchor fragments, pharmacophore points, and mandatory pocket-ligand interactions. To enable this evaluation, we introduce 3D-Fit - a token-efficient benchmarking strategy for assessing LLM performance on multi-conditioned spatial molecule generation. Our findings reveal a clear pattern in LLM spatial capabilities: while they still lag behind state-of-the-art approaches, they are promising and can handle multiple spatial constraints simultaneously, enabling scaling to heterogeneous setups.