arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MolSight: 一种用于统一化学图像理解的图感知视觉语言模型

MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding

Wenda Wang, Yihan Tong, Yuwei Hu, Xuchen Pan, Zhewei Wei, Yaliang Li, Bolin Ding

arXiv 2607.01982首次发表:更新:

发表机构

Renmin University of China(中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MolSight框架,通过分子拓扑模块和分子接地模块增强视觉语言模型对分子图像的结构理解,在多项化学视觉理解任务中显著优于现有模型。

AI 中文摘要

使用分子大语言模型(LLMs)作为理解分子结构和功能的统一框架,在分子设计和药物发现等任务中正成为新趋势。然而,这些模型难以完全捕捉分子结构的视觉表示,限制了其潜力。尽管现有的分子视觉语言模型(VLMs)显示出前景,但它们在结构对齐方面仍面临挑战,且缺乏准确分子理解所需的拓扑建模。为解决这一问题,我们提出了MolSight,一种图感知的视觉语言模型框架,旨在增强VLMs对分子图像的理解。MolSight集成了分子拓扑模块,将化学键邻接信息注入视觉标记,以及分子接地模块,将视觉特征与化学符号语义对齐。我们的实验表明,MolSight在多个化学视觉理解任务中显著优于现有的VLMs、分子LLMs和专用工具,达到了分子图像推理的新水平。

英文摘要

Using molecular large language models (LLMs) as a unified framework for understanding molecular structures and functions is emerging as a new trend in tasks such as molecular design and drug discovery. However, these models struggle to fully capture the visual representation of molecular structures, limiting their potential. While existing molecular vision-language models (VLMs) show promise, they still face challenges in structural alignment and lack the necessary topological modeling for accurate molecular understanding. To address this, we propose MolSight, a graph-aware vision-language model framework designed to enhance the understanding of molecular images by VLMs. MolSight integrates a Molecular Topology Module to inject chemical-bond adjacency information into vision tokens, and a Molecular Grounding Module to align visual features with chemical symbolic semantics. Our experiments demonstrate that MolSight significantly outperforms existing VLMs, molecular LLMs, and task-specific models across multiple chemical visual understanding tasks, achieving a new level of molecular image reasoning in complex chemical scenarios.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑