GLaQ:基于视觉证据的潜在查询定位用于多模态推理
GLaQ: Grounding Latent Queries in Visual Evidence for Multimodal Reasoning
浏览论文内容
中文总结 AI 辅助
该研究提出GLaQ框架,用定位的上下文条件查询替代顺序潜在展开,经训练后在五个视觉理解基准上优于基础模型及其他视觉潜在方法,可高效恢复局部视觉证据。
中文摘要 AI 辅助
思维链推理大幅提升了多模态大语言模型的问题解决能力,然而细粒度视觉证据在基于文本的推理步骤中难以保留和复用。为解决这一局限,借助工具增强的带图像思维方法通过重新访问或操作图像在外部保持视觉访问,但需要预定义工具和额外的推理时处理。作为一种内部替代方案,连续视觉潜在推理将中间计算保留在隐藏状态中,但其普遍采用的自回归结构使得每个潜在状态依赖于前序状态,导致后续状态可能重复潜在序列中已存在的信息,而非捕获互补的视觉细节。我们提出GLaQ,一种基于定位的潜在查询框架,用固定的、基于原始视觉标记的上下文条件查询替代顺序潜在展开。这些定位查询被重新注入以生成答案,提供对源视觉证据的直接且协调的访问。我们采用局部视图监督训练GLaQ,随后在任务级奖励下进行强化学习。在五个细粒度视觉理解与感知基准上,GLaQ-7B较其基础模型提升了5.99--9.66%,且领先所有对比的视觉潜在方法,表明直接的查询到图像定位可从完整图像中恢复局部证据,无需外部视觉操作或自回归潜在展开。
英文摘要
Chain-of-thought reasoning has substantially improved the problem-solving capabilities of multimodal large language models. Fine-grained visual evidence, however, remains difficult to preserve and reuse across text-based reasoning steps. To address this limitation, tool-augmented thinking-with-images methods maintain visual access externally by revisiting or manipulating the image, but require predefined tools and additional inference-time processing. As an internal alternative, continuous visual latent reasoning retains intermediate computation in hidden states. However, its prevailing autoregressive construction makes each latent state depend on its predecessors, so later states may repeat information already present in the latent sequence rather than capture complementary visual details. We introduce GLaQ, a grounded latent-query framework that replaces sequential latent rollout with a fixed set of context-conditioned queries grounded in the original visual tokens. The grounded queries are reinjected for answer generation, providing direct and coordinated access to source visual evidence. We train GLaQ with localized-view supervision followed by reinforcement learning under task-level rewards. Across five benchmarks for fine-grained visual understanding and perception, GLaQ-7B gains 5.99--9.66\% over its base model and leads all compared visual latent methods, suggesting that direct query-to-image grounding can recover localized evidence from the full image without external visual operations or autoregressive latent rollouts.