视觉并行搜索:通过并行瓦片检查与自适应缩放学习搜索高分辨率图像
Visual Parallel Search: Learning to Search High-Resolution Images with Parallel Tile Inspection and Adaptive Zoom
浏览论文内容
中文总结 AI 辅助
针对高分辨率视觉问答中局部证据获取难题,提出VPS并行搜索框架,通过并行瓦片检查与自适应缩放提升准确率,并引入监督与强化学习优化搜索行为。
中文摘要 AI 辅助
高分辨率视觉问答常常因多模态模型未能获取回答所需的小而空间局部的证据而失败。顺序缩放可以恢复细节,但它要求主模型在获得可靠概览之前就选择区域。我们引入了VPS,一种视觉并行搜索框架,其中主智能体首先调用grid_search,与基于问题的子智能体并行检查图像瓦片,然后自适应地调用zoom_in。视觉并行搜索在15个同模型比较中的14个中,相对于专用仅缩放搜索提高了平均准确率,增益高达8.0个百分点,尤其对较小的主模型提升显著。ZoomBench在每种测试规模下保持约3.2个百分点的增益。我们进一步开发了一个带有无提示验证的监督流水线和配对角色特定GRPO替代,用于学习控制器和瓦片读取器角色。SFT在所有五个基准分割上提高了观察到的准确率,包括在HR-Bench 4K上获得4.17个百分点的增益。角色特定RL进一步重塑了搜索行为:仅主RL将内部四响应评估中的平均工具使用从2.65降至2.11,pass@1相似,而外部准确率变化不一。联合训练揭示了局部证据读取与全局搜索控制之间的不对称性。总之,这些结果支持VPS作为有效的推理时脚手架和可训练分解,用于视觉证据获取。
英文摘要
High-resolution visual question answering often fails because a multimodal model does not acquire the small, spatially localized evidence needed to answer a question. Sequential zooming can recover detail, but it asks the main model to choose a region before obtaining a reliable overview. We introduce VPS, a visual parallel-search framework in which a main agent first invokes grid_search to inspect image tiles in parallel with question-conditioned sub-agents, and then adaptively invokes zoom_in on a precise or merged region. The same interface supports both training-free inference and post-training of the main and sub-agents. Across five benchmark splits and three model sizes, VPS improves mean accuracy over dedicated zoom-only search in 14 of 15 same-model comparisons, with gains up to 8.0 points and especially strong improvements for smaller main models. ZoomBench retains an approximately 3.2-point gain at every tested size. We further develop a supervision pipeline with hint-free verification and a paired role-specific GRPO surrogate for learning the controller and tile-reader roles. SFT improves observed accuracy on all five benchmark splits, including a 4.17-point gain on HR-Bench 4K. Role-specific RL further reshapes search behavior: main-only RL reduces mean tool use from 2.65 to 2.11 with similar pass@1 in an internal four-response evaluation, while external accuracy changes are mixed. Joint training reveals an asymmetry between local evidence reading and global search control. Together, these results support VPS as an effective inference-time scaffold and a trainable decomposition for visual evidence acquisition.
发表机构
- The University of Hong Kong(香港大学)
- Huawei Research(华为研究院)
- City University of Hong Kong(香港城市大学)
- Huazhong University of Science and Technology(华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。