arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MolEmb:多模态大语言模型可作为强大的分子嵌入模型

MolEmb: Multimodal Large Language Models Can Be Strong Molecular Embedding Models

Xinjian Zhao, Xiangru Jian, Yaoyao Xu, Xiaozhuang Song, Wei Pang, Lei Bai, Tianshu Yu

arXiv 2608.23646首次发表:更新:

AI 中文总结

本文提出MolEmb框架,利用多模态大语言模型构建通用分子嵌入模型,在分子性质预测上具竞争力,还提出MolCAR基准验证上下文感知分子嵌入的特性。

AI 中文摘要

分子嵌入模型可作为计算化学与药物发现的基础基础设施,其可复用的向量表示支持性质预测、虚拟筛选与检索。大多数分子编码器是围绕单一分子视图构建的专用模型,生成无条件向量,且无用于调整表示的语言接口。本文探究原生处理图像、文本与符号输入的多模态大语言模型(MLLMs)是否可作为通用分子嵌入模型,生成同时基于分子特征与自然语言语义上下文的嵌入。本文提出轻量级框架MolEmb,通过双向对比目标将分子特征与文本描述在共享嵌入空间中对齐,以适配MLLMs。所得嵌入模型在分子性质预测上具有竞争力,且支持同一空间内的跨模态分子-文本检索。本文进一步提出用于上下文感知检索的诊断基准MolCAR,发现上下文感知分子嵌入主要是监督的数据属性。这些结果表明,MLLMs不仅是化学助手或生成器,也是通用分子嵌入模型的可行且可扩展的路径。

英文摘要

Molecular embedding models can serve as foundational infrastructure for computational chemistry and drug discovery, where reusable vector representations support property prediction, virtual screening, and retrieval. Most molecular encoders are specialist models built around a single molecular view, producing unconditional vectors with no language interface for varying the representation. We ask whether multimodal large language models (MLLMs), which natively process images, text, and symbolic inputs, can instead serve as \emph{general molecular embedding models} that produce embeddings conditioned on both a molecular profile and a natural-language semantic context. We introduce \textbf{MolEmb}, a lightweight framework that adapts MLLMs by aligning molecular profiles with textual descriptions in a shared embedding space using a bidirectional contrastive objective. The resulting embedding model is competitive on molecular property prediction and supports cross-modal molecule--text retrieval in the same space. We further introduce \textbf{MolCAR}, a diagnostic benchmark for context-aware retrieval, and find that context-aware molecular embedding is primarily a data property of the supervision. These results suggest that MLLMs are not merely chemistry assistants or generators, but a viable and extensible route to general molecular embedding models.

CommentsPresented at the 3rd Workshop on Multi-modal Foundation Models and Large Language Models for Life Sciences (FM4LS), ICML 2026. Non-archival workshop

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑