arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27830cs.CV

一次思考足矣:面向高分辨率视觉问答的中间层证据路由

Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA

Zhongkuan Mao, Xianjie Liu, Tianyu Meng, Yidong Wang, Wenzhuo Zhao, Ronghao Xian, Yao Jiang, Fei Shen, Junfeng Fang, Yong Dai, Yi Zhang, Keren Fu

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对高分辨率视觉问答提出无需训练的Thinking-Once证据路由方法,通过利用中间层留存的细粒度证据,在多个基准上提升性能并降低内存与推理时间。

中文摘要 AI 辅助

高分辨率视觉问答(HR-VQA)常被视为证据获取不足的问题,表现不佳的多模态大语言模型需通过裁剪、重编码或多轮搜索再次检查图像。本文指出该观点不完整:很多情况下,细粒度证据已在视觉编码中留存,在中间层路由窗口内可被识别且具影响力,但在答案生成前被稀释。我们提出Thinking-Once,这是一种无需训练、单次视觉传递的证据路由方法,可在该窗口重构基于问题的注意力,保留核心实体 token 和紧凑背景上下文,并将此证据路由至后续层,无需额外视觉编码。在五个基础模型上,Thinking-Once始终优于或匹配对应基础设置,使V$^*$Bench、HRBench-4K、HRBench-8K的平均分数分别提升+3.1、+3.0、+2.7个点,同时平均峰值内存减少约4GB。在Qwen2.5-VL-7B上,它使三个基准分别提升+9.9、+4.6、+5.5个点,跨基准均值从72.5升至79.1;在ZwZ-8B基础模型上,Thinking-Once达到82.7的均值。与11个开源HR-VQA基线相比,它在三个基准平均上取得最佳或并列最佳分数及最佳总体均值;例如,与DeepScan相比,它将V$^*$Bench推理时间减少97.2%,同时跨基准均值从77.8提升至79.1。这些结果表明,HR-VQA可通过路由已编码证据而非重复获取新视觉输入来改进,代码见附录。

英文摘要

High-resolution visual question answering (HR-VQA) is often treated as a problem of insufficient evidence acquisition, where failing multimodal large language models must inspect images again through cropping, re-encoding, or multi-round search. We show that this view is incomplete: in many cases, fine-grained evidence has already survived visual encoding and become identifiable and influential within an intermediate-layer routing window, but is later diluted before answer generation. We propose Thinking-Once, a \textbf{training-free, single-visual-pass} evidence-routing method that reconstructs question-conditioned attention at this window, preserves core entity tokens and compact background context, and routes this evidence to later layers without extra visual encoding. Across five base models, Thinking-Once consistently improves or matches the corresponding base setting, increasing the average scores on V$^*$Bench, HRBench-4K, and HRBench-8K by \textit{+3.1}, \textit{+3.0}, and \textit{+2.7} points while reducing the average peak memory by about 4,GB. On Qwen2.5-VL-7B, it improves the three benchmarks by \textit{+9.9}, \textit{+4.6}, and \textit{+5.5} points, raising the cross-benchmark mean from 72.5 to 79.1. With the ZwZ-8B base model, Thinking-Once reaches a mean score of 82.7. Against 11 open-source HR-VQA baselines, it obtains the best or tied-best score on all three benchmark averages and the best overall mean; for example, compared with DeepScan, it reduces V$^*$Bench inference time by \textbf{97.2\%} while improving the cross-benchmark mean from 77.8 to 79.1. These results show that HR-VQA can be improved by routing already encoded evidence rather than repeatedly acquiring new visual inputs. Code is available in the appendix.

↑