检索图像作为视觉思维:面向开放与封闭差距的无训练多模态上下文学习
Retrieved Images as Visual Thought: Training-Free Multimodal In-Context Learning for the Open-vs-Closed Gap
浏览论文内容
中文总结 AI 辅助
提出无训练框架ReVisIT,通过检索图像-标签对作为视觉思维单元,结合结构化类定义、多模态检索和交替注入示例,在多个基准上显著提升性能,弥合开放与封闭任务差距。
中文摘要 AI 辅助
最近关于“用图像思考”的工作使视觉成为推理的动态部分,但通过生成实现:模型调用外部工具、合成代码或想象新图像,每种方式都需付出工具协议、脆弱代码或昂贵训练管道的代价。第四条路径无需生成任何内容,通过检索带标签的示例图像并对其进行推理,使视觉变得动态,尽管无需训练,但仍未得到充分探索。我们提出ReVisIT,一个无训练框架,通过将每个检索到的图像-标签对视为视觉思维单元来实现这一基于检索的路径。ReVisIT结合了结构化类定义、每查询多模态示例检索、在联合多属性解码前交替注入用户/助手示例,并能优雅地降级到任务允许的任何组件。在VL-ICL Bench Fast Open MiniImageNet上,使用ReVisIT的Qwen3-VL-30B-A3B在4-shot下达到98.5%,与72B LLaVA-OneVision SOTA(98.7%)在统计上无显著差异,参数约为其1/2.4,而相同骨干网络在没有该框架时仅处于随机水平。仅交替层就在自由形式概念归纳(Bongard-OpenWorld)上为GPT-4.1增加了26.1个点,完整堆栈在MAAC-Bench(一个新的许可证清洁的27类、5属性基准)上跨三个骨干网络实现了4-6个点的宏观增益,通过配对bootstrap在策展派生属性上显著。组件分析表明,检索加交替是通用杠杆,而结构化定义是需求自适应的,并且83%的检索增益来自检索质量而非示例的存在。MAAC-Bench附带一个基于量规的LLM验证协议发布,该协议取代了作者对主观属性的抽查。
英文摘要
Recent work on Thinking with Images makes vision a dynamic part of reasoning, but does so through generation: the model invokes external tools, synthesizes code, or imagines new imagery, each at the cost of a tool protocol, brittle code, or an expensive training pipeline. A fourth route makes vision dynamic without generating anything, by retrieving labeled exemplar images and reasoning over them, yet it remains underexplored despite being train-free. We present ReVisIT, a train-free framework that realizes this retrieval-based route by treating each retrieved image-label pair as a unit of visual thought. ReVisIT combines structured class definitions, per-query multimodal retrieval of exemplars, and alternating user/assistant injection of those exemplars before joint multi-attribute decoding, and degrades gracefully to whichever components a task admits. On VL-ICL Bench Fast Open MiniImageNet, Qwen3-VL-30B-A3B with ReVisIT reaches 98.5% at 4-shot, statistically indistinguishable from the 72B LLaVA-OneVision SOTA (98.7%) on this near-saturated task at about 1/2.4 the parameters, while the same backbone without the scaffold sits at chance. The turns layer alone adds 26.1 points to GPT-4.1 on free-form concept induction (Bongard-OpenWorld), and the full stack yields a 4-6 point macro gain across three backbones on MAAC-Bench, a new license-clean 27-class, 5-attribute benchmark, significant by paired bootstrap on the curator-derived attributes. Component analysis shows that retrieval-plus-turns is the universal lever while structured definitions are need-adaptive, and that 83% of the retrieval gain comes from retrieval quality rather than from the presence of exemplars. MAAC-Bench is released with a rubric-grounded LLM verification protocol that replaces author spot-check on subjective attributes.