LiteSearch-VL:通过轨迹蒸馏和合成步级DPO实现的小型多模态搜索智能体
LiteSearch-VL: Small Multimodal Search Agents via Trajectory Distillation and Synthetic Step-DPO
- Microsoft AI(微软人工智能)
- Ohio State University(俄亥俄州立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出LiteSearch-VL方案,通过轨迹蒸馏和合成步级DPO,在单节点预算下将多模态搜索智能体能力迁移至Qwen3-VL-2B等小型模型,使其在多模态视觉问答任务上取得接近4B模型的性能,明确小型多模态智能体的下一瓶颈是答案验证而非搜索深度。
AI中文摘要:
多模态搜索智能体通过交替进行图像理解、网络检索、工具使用和证据合成来回答视觉问题。目前已存在强大的系统,但它们属于两种高成本范畴:一类是GPT-5、Gemini等专有前沿模型,另一类是用大量智能体数据和强化学习训练的大型开放视觉-语言主干模型。本文提出了一个不同的问题:当将发布的智能体轨迹在单节点预算下蒸馏到小得多的主干模型中时,究竟会传递什么?我们用LiteSearch-VL研究这一问题,它是针对Qwen3-VL-2B和Qwen3-VL-4B的低计算量方案,仅使用发布的OpenSearch-VL轨迹、参数高效的LoRA适配器,以及合成步级偏好:针对五种局部失败模式(过早回答、工具错误、查询薄弱、查询重复、忽略图像)的GPT-5生成的难负例上的DPO。在SimpleVQA、FVQA、LiveVQA和VDR-Bench-testmini上的12400个GPT-5评估的 rollout 中,主导效应是行为层面的而非均匀的准确率提升:全轨迹监督微调传递了智能体契约,使2B模型几乎从不输出可用答案(1240次rollout中1237次为no_answer)提升至28.4%的宏Pass@1,匹配或略超过现成的4B基础模型(25.6%)。合成偏好学习和紧凑工具蒸馏作为细化而非相变(最佳4B配置:30.8%宏Pass@1)。最后,受控VDR步骤预算消融显示,额外搜索轮次将弃权(不执行)转化为实体错误而非正确答案,确定答案验证而非搜索深度是小型多模态智能体的下一个瓶颈。
英文摘要:
Multimodal search agents answer visual questions by interleaving image understanding, web retrieval, tool use, and evidence synthesis. Strong systems exist, but in two expensive regimes: proprietary frontier models such as GPT-5 and Gemini, or large open vision-language backbones trained with substantial agentic data and reinforcement learning. We ask a different question: when released agent trajectories are distilled into much smaller backbones under a single-node budget, what is actually transferred? We study this with LiteSearch-VL, a low-compute recipe for Qwen3-VL-2B and Qwen3-VL-4B that uses only released OpenSearch-VL trajectories, parameter-efficient LoRA adapters, and synthetic step-level preferences: DPO on GPT-5-generated hard negatives targeting five local failure modes (premature answer, wrong tool, weak query, repeated query, ignored image). Across 12,400 GPT-5-judged rollouts on SimpleVQA, FVQA, LiveVQA, and VDR-Bench-testmini, the dominant effect is behavioral rather than a uniform accuracy lift: full-trajectory supervised fine-tuning transfers the agent contract, taking the 2B model from almost never emitting a usable answer (1,237/1,240 no_answer rollouts) to 28.4% macro Pass@1, matching or slightly exceeding the off-the-shelf 4B base (25.6%). Synthetic preference learning and compact tool distillation act as refinements rather than phase transitions (best 4B configuration: 30.8% macro Pass@1). Finally, a controlled VDR step-budget ablation shows that extra search turns convert abstentions into wrong_entity errors rather than correct answers, identifying answer verification, not search depth, as the next bottleneck for small multimodal agents.