VisLens:面向多模态大语言模型的单次可解释视觉搜索方法
VisLens: Single-Pass Interpretable Visual Search for Multimodal LLMs
浏览论文内容
中文总结 AI 辅助
该研究针对多模态大语言模型的细粒度视觉搜索问题,提出VisLens方法,通过单次前向传播实现可解释的视觉搜索,性能优于或相当现有方法且推理速度大幅提升。
中文摘要 AI 辅助
多模态大语言模型(MLLMs)在细粒度视觉搜索任务中表现不佳,该任务旨在从高分辨率图像中定位小型或罕见物体。现有解决方案分为两类:(1)基于注意力或置信度分数的无训练方法准确率高但速度慢,因为每个示例需要多次查询MLLM;(2)通过强化学习(RL)训练的工具使用模型推理速度更快,但不透明,因为其工具调用无法控制且难以解释。为解决这一问题,我们提出VisLens(Visual Focus via Logit Lens,基于logit透镜的视觉聚焦),这是一种基于logit透镜构建的视觉搜索方法,通过将隐藏状态投影到LLM头部来解码其中包含的语义。VisLens还使用轻量级调优透镜将早期隐藏状态映射到最终隐藏状态空间,从而可以从早期层中读取视觉令牌。这些令牌与查询中的目标词匹配,生成相关区域的裁剪图像,与原始图像一起反馈以生成最终答案。整个过程从解码到最终答案在单次前向传播中完成,无需重复查询。VisLens的性能与现有基准相当或更优,同时具有显著的延迟优势,运行速度比Thyme快8.5至9.9倍,比无训练的多轮搜索方法快达22.2倍。
英文摘要
Multimodal large language models (MLLMs) struggle with fine-grained Visual Search, the task of locating small or rare objects in high-resolution images. Existing remedies fall into two families: (1) Training-free methods based on attention or confidence scores are accurate but slow, since they require multiple MLLM queries per example. (2) Reinforcement Learning (RL) trained tool-use models are faster at inference but opaque, since their tool calls remain uncontrollable and hard to interpret. To overcome this, we propose \emph{VisLens} (Visual Focus via Logit Lens), a Visual Search method built on the logit lens, which decodes the semantics held in a hidden state by projecting it through the LLM head. VisLens further uses a lightweight tuned-lens that maps early hidden states into the final hidden state space, so visual tokens can be read out from early layers. These tokens are matched to target words in the query to generate a crop of the relevant region, which is fed back in alongside the original image to produce the final answer. The whole process, from decoding to the final answer, completes in a single forward pass without repeated queries. VisLens matches or exceeds prior baselines while delivering a substantial latency advantage, running $8.5$--$9.9\times$ faster than Thyme and up to $22.2\times$ faster than training-free multi-pass search methods.
发表机构
- Technical University of Munich(慕尼黑工业大学)
- Helmholtz Munich(慕尼黑亥姆霍兹中心)
- Munich Center for Machine Learning(慕尼黑机器学习中心)
机构由 AI 辅助整理,请以论文原文为准。