arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在视觉空间中通过多步嵌入检索学习路由

Learning to Route in Visual Space via Multi-Step Embedding Retrieval

Tianyu Chen, Mingyuan Zhou, Jiaxing Wu

arXiv 2609.38743首次发表:更新:

发表机构

The University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出VHOP框架与VHOP-Router训练流程,将标准嵌入模型转化为多步检索器,在视觉潜在空间中单次调用检索图像链,显著提升检索性能与智能体搜索成功率,并大幅降低令牌与API负载。

AI 中文摘要

LLM智能体依赖检索工具访问外部知识,然而视觉智能体搜索仍严重受限于标准的单步检索器。在当前流程中,智能体必须为每个中间步骤发出文本查询,当视觉线索难以描述或检索器未能在其顶部结果中呈现必要的中间证据时,这一过程会陷入困境。我们假设将跨越整个嵌入空间的多步导航直接卸载给检索工具能够解决这一性能瓶颈。为系统性地研究这一问题,我们引入了VHOP,一个灵活的数据生成框架和基准测试,包含五个核心难度级别,同时测试视觉匹配和搜索规划能力。利用该框架,我们开发了VHOP-Router,一个端到端训练流程——结合监督微调、在线模仿学习和强化学习——将标准嵌入模型转变为自回归多步检索器。VHOP-Router直接在视觉潜在空间中操作,通过单次工具调用检索链接的图像链,无需智能体制定中间文本查询。实验表明,VHOP-Router将检索性能从不足5%提升至76.3%。在智能体搜索中,它将任务成功率提高了52.7%,并将平均令牌长度从1886减少至728,降低了61%,而升级智能体仅带来3.7%的提升。与智能体每步检索前50个结果的强基线相比,VHOP-Router在保持优越性能的同时,将上下文图像减少了23倍,并将累计API负载削减了35倍。该模型还能稳健地泛化到未见过的难度级别和现实测试集。最终,VHOP和VHOP-Router为视觉智能体搜索提供了一种高效且有效的解决方案,且完全保留原生LLM能力。

英文摘要

LLM agents rely on retrieval tools to access external knowledge, yet visual agentic search remains severely bottlenecked by standard single-step retrievers. In current pipelines, the agent must issue text queries for every intermediate step, struggling when visual clues are difficult to describe or when the retriever fails to surface necessary intermediate evidence within its top results. We hypothesize that offloading multi-step navigation across the entire embedding space directly to the retrieval tool resolves this performance bottleneck. To study this systematically, we introduce VHOP, a flexible data generation framework and benchmark with five core difficulty levels testing both visual matching and search planning. Using this framework, we develop VHOP-Router, an end-to-end training pipeline---combining supervised fine-tuning, online imitation learning, and reinforcement learning---that transforms a standard embedding model into an autoregressive multi-step retriever. Operating directly in the visual latent space, VHOP-Router retrieves linked image chains in a single tool call without requiring the agent to formulate intermediate text queries. Experiments show VHOP-Router boosts retrieval performance from under 5\% to 76.3\%. In agentic search, it improves task success rates by 52.7\% and reduces the average token length by 61\% from 1886 to 728, whereas upgrading the agent yields only a 3.7\% gain. Compared to a strong baseline where the agent retrieves the top 50 results per step, VHOP-Router maintains superior performance while reducing in-context images by $23\times$ and cutting the cumulative API payload by $35\times$. The models also generalize robustly to unseen difficulty levels and realistic test sets. Ultimately, VHOP and VHOP-Router provide an efficient and effective solution for visual agentic search that leaves native LLM capabilities entirely intact.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑