CoMPASS:通过自适应大小模型协同实现协同分子性质预测
CoMPASS: Collaborative Molecular Property Prediction via Adaptive Small-Large Model Synergy
浏览论文内容
中文总结 AI 辅助
CoMPASS是一种用于大小模型协同的检索校准框架,通过结合GAT与LLM的优势,在分子性质预测任务中提升了可修正不确定性区域的性能并限制了LLM的不当干预。
中文摘要 AI 辅助
准确的分子性质预测需要统计可靠性与化学推理能力。图神经网络(GAT)可直接在标记实验数据上校准,但受限于训练数据的覆盖范围;大语言模型(LLM)能够比对分子证据并阐述化学原理,但作为独立定量预测器不可靠。核心挑战在于确定LLM应在何时影响校准模型及影响程度。本文提出CoMPASS,一种用于大小模型协同的检索校准框架:保留图注意力网络(GAT)作为预测锚点,检索局部相关的训练分子,为LLM提供基于注意力的证据,通过感知一致性的门控将LLM的建议转换为有界修正。在6个分类和2个回归基准上,CoMPASS在可修正不确定性区域提升了GAT锚点的性能,同时在高置信度场景限制LLM的干预。消融实验表明,性能提升源于经验证校准的检索与有界融合,而非仅提示工程。这些结果表明,生成式推理应通过基于证据的可控修正来增强校准预测,而非直接替换输出。代码可在此https URL获取。
英文摘要
Accurate molecular property prediction requires both statistical reliability and chemical reasoning. Graph neural networks can be calibrated directly on labeled assays but remain limited by the coverage of their training data. Large language models (LLMs) can compare molecular evidence and articulate chemical rationales, yet are unreliable as standalone quantitative predictors. The central challenge is therefore to determine when an LLM should influence a calibrated model and by how much. Here we present CoMPASS, a retrieval-calibrated framework for small-large model collaboration. CoMPASS retains a graph attention network (GAT) as the predictive anchor, retrieves locally relevant training molecules, provides attention-grounded evidence to an LLM, and converts its proposal into a bounded correction through an agreement-aware gate. Across six classification and two regression benchmarks, CoMPASS improves the GAT anchor in regions of correctable uncertainty while limiting LLM intervention in high-confidence regimes. Ablations show that the gains arise from validation-calibrated retrieval and bounded fusion rather than prompting alone. These results suggest that generative reasoning should augment calibrated prediction through evidence-grounded, controlled corrections rather than direct output replacement. Code is available at https://github.com/littlepeachs/CoMPASS.
发表机构
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。