arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DeepVoyager-VL:为长视野多模态智能体激励视觉在环搜索

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents

Huanyao Zhang, Jiepeng Zhou, Runhao Zhao, Yanzhe Shan, Jiaoyang Chen, Bowen Zhou, Bo Li, Fang Wang, Jialong Wu, Zhengwei Tao, Lang Mei, Xiaohan Yu, Liyan Liu, Chong Chen, Wentao Zhang

arXiv 2608.01827首次发表:更新:

发表机构

PKU; HKUST(GZ); NUDT; OUC; HITSZ; Huawei Cloud BU(北京大学; 香港科技大学(广州); 中国人民解放军国防科技大学; 中国海洋大学; 哈尔滨工业大学(深圳); 华为云业务部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有多模态深度搜索忽视视觉中间推理作用的局限,提出DeepVoyager-VL框架,构建多模态事件图合成数据,设计智能体框架实现主动视觉获取,经10个基准实验验证其有效性。

AI 中文摘要

多模态大语言模型(MLLMs)已在视觉理解与推理方面取得进展,但其静态参数知识限制了解决知识密集型及动态演化开放世界问题的能力。为突破这一局限,多模态深度搜索作为开放世界信息获取的关键方向应运而生,正从单轮事实检索向视觉证据引导的长视野多轮搜索演进。然而,现有方法通常将视觉限制在输入或答案阶段,忽视其在中间推理中的作用,且缺乏针对长视野交互的定制设计。因此,视觉证据极少驱动持续检索,限制了交互深度与推理跨度。为解决这些局限,我们提出DeepVoyager-VL,一种用于视觉在环搜索的长视野多模态深度搜索框架。具体而言,我们构建多模态事件图以驱动数据合成,生成具有中间视觉依赖与长推理链的问题;接着设计智能体框架以实现主动视觉获取与按需图像加载;最后在合成数据上对模型进行微调,不使用强化学习。在10个多模态搜索基准上开展的大量实验证明了我们方法的有效性。

英文摘要

Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive and dynamically evolving open-world problems. To move beyond this limitation, multimodal deep search has emerged as a key direction for open-world information access, evolving from single-turn factual retrieval toward long-horizon, multi-turn search guided by visual evidence. However, existing methods typically confine vision to the input or answer stage, overlooking its role in intermediate reasoning, and lack designs tailored to long-horizon interaction. Consequently, visual evidence rarely drives continued retrieval, constraining both interaction depth and reasoning span. To address these limitations, we propose DeepVoyager-VL, a long-horizon multimodal deep-search framework for vision-in-the-loop search. Specifically, we construct a multimodal event graph to drive data synthesis, yielding problems with intermediate visual dependencies and long reasoning chains. We then design an agent framework for active visual acquisition and on-demand image loading. Finally, we fine-tune models on the synthesized data without reinforcement learning. Extensive experiments across ten multimodal search benchmarks demonstrate the effectiveness of our method.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑