NeuronEye:查询引导的视觉概念激活用于视觉-语言推理
NeuronEye: Query-Guided Visual Concept Activation for Vision-Language Reasoning
浏览论文内容
中文总结 AI 辅助
NeuronEye通过构建稀疏概念级神经元词汇表并利用语言查询激活相关视觉概念,在冻结VLM上单次前向传播提升视觉-语言推理性能,如CV-Bench提升3.1点。
中文摘要 AI 辅助
当前的视觉-语言模型(VLM)将视觉信息编码在密集的隐藏状态中,其中物体身份、空间布局和局部属性隐式纠缠而非显式解耦,限制了它们隔离和调节给定语言查询所需的具体视觉证据的能力。受生物视觉中稀疏群体编码和自上而下调节的启发,我们引入了NeuronEye,一个即插即用框架,从VLM中间表示构建稀疏的概念级神经元词汇表,并在推理过程中选择性激活与查询相关的视觉概念。NeuronEye将视觉令牌状态分解为按概念级聚类组织的过完备稀疏基,利用语言查询激活相关聚类并定位所选概念表达的补丁,然后将聚焦的证据注入回视觉令牌。一种互补的抑制机制衰减主导感知方向以保留较弱但相关的线索。所有操作在冻结的VLM骨干网络上的单次前向传播中运行。在Qwen2.5-VL-7B上,NeuronEye将CV-Bench整体准确率提升+3.1,其中Distance提升+9.5,并将BLINK多视图提升+8.3,在LLaVA-1.6-7B上也有类似趋势。这些结果表明,稀疏神经元词汇表不仅可以作为事后可解释性工具,还可以作为概念级视觉推理的主动接口。
英文摘要
Current vision-language models (VLMs) encode visual information in dense hidden states where object identity, spatial layout, and local attributes are implicitly entangled rather than explicitly disentangled, limiting their ability to isolate and modulate the specific visual evidence required by a given language query. Inspired by sparse population coding and top-down modulation in biological vision, we introduce NeuronEye, a plug-in framework that constructs a sparse, concept-level neuron vocabulary from intermediate VLM representations and selectively activates query-relevant visual concepts during inference. NeuronEye decomposes vision-token states into an overcomplete sparse basis organized by concept-level clusters, uses the language query to activate relevant clusters and localize the patches where selected concepts are expressed, and injects the focused evidence back into vision tokens. A complementary suppression mechanism attenuates dominant perceptual directions to preserve weaker but relevant cues. All operations run in a single forward pass over a frozen VLM backbone. On Qwen2.5-VL-7B, NeuronEye raises CV-Bench overall accuracy by +3.1 with gains of +9.5 on Distance, and improves BLINK Multi-view by +8.3, with similar trends on LLaVA-1.6-7B. These results suggest that sparse neuron vocabularies can serve not only as post-hoc interpretability tools but also as active interfaces for concept-level visual reasoning.
发表机构
- New York University(纽约大学)
- New Jersey Institute of Technology(新泽西理工学院)
机构由 AI 辅助整理,请以论文原文为准。