一次思考足矣:面向高分辨率视觉问答的中间层证据路由
Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA
浏览论文内容
中文总结 AI 辅助
本文针对高分辨率视觉问答提出无需训练的Thinking-Once证据路由方法,通过利用中间层留存的细粒度证据,在多个基准上提升性能并降低内存与推理时间。
中文摘要 AI 辅助
高分辨率视觉问答(HR-VQA)常被视为证据获取不足的问题,表现不佳的多模态大语言模型需通过裁剪、重编码或多轮搜索再次检查图像。本文指出该观点不完整:很多情况下,细粒度证据已在视觉编码中留存,在中间层路由窗口内可被识别且具影响力,但在答案生成前被稀释。我们提出Thinking-Once,这是一种无需训练、单次视觉传递的证据路由方法,可在该窗口重构基于问题的注意力,保留核心实体 token 和紧凑背景上下文,并将此证据路由至后续层,无需额外视觉编码。在五个基础模型上,Thinking-Once始终优于或匹配对应基础设置,使V$^*$Bench、HRBench-4K、HRBench-8K的平均分数分别提升+3.1、+3.0、+2.7个点,同时平均峰值内存减少约4GB。在Qwen2.5-VL-7B上,它使三个基准分别提升+9.9、+4.6、+5.5个点,跨基准均值从72.5升至79.1;在ZwZ-8B基础模型上,Thinking-Once达到82.7的均值。与11个开源HR-VQA基线相比,它在三个基准平均上取得最佳或并列最佳分数及最佳总体均值;例如,与DeepScan相比,它将V$^*$Bench推理时间减少97.2%,同时跨基准均值从77.8提升至79.1。这些结果表明,HR-VQA可通过路由已编码证据而非重复获取新视觉输入来改进,代码见附录。
英文摘要
High-resolution visual question answering (HR-VQA) is often treated as a problem of insufficient evidence acquisition, where failing multimodal large language models must inspect images again through cropping, re-encoding, or multi-round search. We show that this view is incomplete: in many cases, fine-grained evidence has already survived visual encoding and become identifiable and influential within an intermediate-layer routing window, but is later diluted before answer generation. We propose Thinking-Once, a \textbf{training-free, single-visual-pass} evidence-routing method that reconstructs question-conditioned attention at this window, preserves core entity tokens and compact background context, and routes this evidence to later layers without extra visual encoding. Across five base models, Thinking-Once consistently improves or matches the corresponding base setting, increasing the average scores on V$^*$Bench, HRBench-4K, and HRBench-8K by \textit{+3.1}, \textit{+3.0}, and \textit{+2.7} points while reducing the average peak memory by about 4,GB. On Qwen2.5-VL-7B, it improves the three benchmarks by \textit{+9.9}, \textit{+4.6}, and \textit{+5.5} points, raising the cross-benchmark mean from 72.5 to 79.1. With the ZwZ-8B base model, Thinking-Once reaches a mean score of 82.7. Against 11 open-source HR-VQA baselines, it obtains the best or tied-best score on all three benchmark averages and the best overall mean; for example, compared with DeepScan, it reduces V$^*$Bench inference time by \textbf{97.2\%} while improving the cross-benchmark mean from 77.8 to 79.1. These results show that HR-VQA can be improved by routing already encoded evidence rather than repeatedly acquiring new visual inputs. Code is available in the appendix.