AI 中文总结
该研究针对现有化学推理模型无法聚焦分子关键区域的问题,提出VLSR框架,采用先定位再推理策略,在分子性质推理任务上实现9.6倍于文本推理基准的吞吐量提升。
AI 中文摘要
局部化学感知与性质推理对于理解分子结构如何决定性质均至关重要。当前基于大语言模型(LLM)的化学推理方法要么接收SMILES/分子图像及局部基序描述,要么直接从分子图像进行推理,两种方法均无法让模型在推理前聚焦于化学上有意义的区域。为解决这一差距,我们提出视觉隐式结构推理(Visual Latent Structural Reasoning,VLSR),这是一种从分子图像联合学习定位与推理的端到端框架,核心是先定位再推理的策略。VLSR首先学习定位分子图像中化学上有意义的区域,随后在紧凑的隐式工作空间中推理这些区域的性质效应,再生成最终答案。在相同推理设置下,该设计的吞吐量比可比的文本推理基准高9.6倍。
英文摘要
Local chemical perception and property reasoning are both essential for understanding how molecular structure determines properties. Current LLM-based chemical reasoning methods either receive SMILES/molecular images together with descriptions of local motifs, or reason directly from molecular images. Neither approach enables the model to focus on chemically meaningful regions before reasoning. To address this gap, we propose Visual Latent Structural Reasoning (VLSR), an end-to-end framework that jointly learns localization and reasoning from molecular images. Central to our approach is a localize-then-reason strategy. VLSR first learns to locate chemically meaningful regions in a molecular image. It then reasons about their property effects in a compact latent workspace before producing the final answer. Under the same inference setup, this design achieves 9.6X higher throughput than a comparable textual-reasoning baseline.