发表机构
Leiden Institute of Advanced Computer Science (LIACS), Leiden University(莱顿大学高级计算机科学研究所(LIACS))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出BrainFocus框架,利用EEG信号引导VLM仅处理相关ROI,在40类基准上提升VQA准确率4.14-9.87pp并减少23.2%-39.5%计算量。
AI 中文摘要
视觉语言模型(VLMs)在视觉问答(VQA)任务上取得了强劲的性能,但当只有图像中的小区域相关时,处理大型杂乱图像的计算成本很高。脑电图(EEG)信号能够捕捉人类对视觉刺激的神经反应,可以提供关于感兴趣区域(ROI)的基于人类的语义线索。然而,EEG引导的视觉类别解码仍不完美,使得直接进行ROI路由不可靠。在这项工作中,我们提出了BrainFocus,一个可靠的EEG引导的高效VLM框架,用于VQA。一个EEG分类器预测目标类别,一个YOLO检测器定位匹配的ROI。仅当两个预测均通过置信度阈值时,VLM才接收裁剪后的ROI;否则,它处理完整图像。为了评估,我们在EEG-ImageNet的基础上构建了一个包含40个类别的基准,该基准包含生成的杂乱图像和真实的目标中心图像,并带有目标ROI标注和600个英文视觉问答对。在Qwen3.5-VL 2B、4B和9B模型上,BrainFocus在杂乱场景上将VQA准确率提高了4.14-9.87个百分点(pp),同时将输入token和总token分别减少了23.2%-39.4%和23.2%-39.3%,端到端浮点运算量(FLOPs)减少了23.2%-39.5%。这些结果表明,即使EEG语义解码不完美,EEG也能引导高效的VLM推理。
英文摘要
Vision-language models (VLMs) achieve strong visual question answering (VQA) performance, but processing large cluttered images is computationally expensive when only a small region is relevant. Electroencephalography (EEG) signals, which capture human neural responses to visual stimuli, can provide a human-derived semantic cue about the region of interest (ROI). However, EEG-guided visual category decoding remains imperfect, making direct ROI routing unreliable. In this work, we propose BrainFocus, a reliable EEG-guided efficient VLM framework for VQA. An EEG classifier predicts a target category, and a YOLO detector localizes the matching ROI. The VLM receives the cropped ROI only when both predictions pass confidence thresholds; otherwise, it processes the full image. For evaluation, we build on EEG-ImageNet to construct a 40-class benchmark comprising generated cluttered images and real object-centric images, with target-ROI annotations and 600 English visual question-answer pairs. Across Qwen3.5-VL 2B, 4B, and 9B models, BrainFocus improves VQA accuracy by 4.14-9.87 percentage points (pp) on cluttered scenes while reducing input tokens and total tokens by 23.2%-39.4% and 23.2%-39.3%, and end-to-end floating-point operations (FLOPs) by 23.2%-39.5%. These results demonstrate that EEG can guide efficient VLM inference even when its semantic decoding is imperfect.