面向视觉原生多模态深度搜索智能体的在策略数据演化
Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents
- Hong Kong University of Science and Technology(香港理工大学)
- Renmin University of China(中国人民大学)
- The Chinese University of Hong Kong(香港中文大学)
- Peking University(北京大学)
- Tsinghua University(清华大学)
- University of Edinburgh(爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出在策略数据演化(ODE)框架,通过图像库引用协议和闭环数据生成器,解决多模态深度搜索中视觉证据不可复用和训练数据静态问题,在8个基准上显著提升性能。
AI中文摘要:
多模态深度搜索要求智能体通过链式搜索、工具使用和对不断变化的文本与视觉上下文的视觉推理来解决开放世界问题。两个瓶颈限制了当前系统。首先,现有的工具使用框架将搜索、浏览或转换返回的图像视为瞬时输出,因此中间视觉证据无法被后续工具重新消费。其次,训练数据通常由固定的整理配方构建,无法跟踪目标智能体不断发展的能力。为应对这些挑战,我们首先引入了一个以图像库引用协议为核心的视觉原生智能体框架,该协议将每个工具返回的图像注册为可寻址引用,使中间视觉证据可被后续工具重用。在此框架之上,在策略数据演化(ODE)运行一个闭环数据生成器,该生成器根据正在训练的策略的 rollout 在每轮中自我改进。这种逐轮改进使得每轮的数据针对当前策略仍需学习的内容。同一框架支持多样化的监督微调数据和策略感知的强化学习数据整理,覆盖目标智能体的完整训练生命周期。在8个多模态深度搜索基准上,ODE 将 Qwen3-VL-8B 智能体的平均得分从24.9%提升至39.0%,在标准智能体工作流设置中超越了 Gemini-2.5 Pro(37.9%)。在30B规模下,ODE 将平均得分从30.6%提升至41.5%。进一步分析验证了图像库重用的有效性,特别是在需要迭代视觉细化的复杂任务上,而 rollout 反馈演化比静态合成产生了更扎实的 SFT 轨迹和更好的策略匹配的 RL 任务。
英文摘要:
Multimodal deep search requires an agent to solve open-world problems by chaining search, tool use, and visual reasoning over evolving textual and visual context. Two bottlenecks limit current systems. First, existing tool-use harnesses treat images returned by search, browsing, or transformation as transient outputs, so intermediate visual evidence cannot be re-consumed by later tools. Second, training data is usually built by fixed curation recipes that cannot track the target agent's evolving capability. To address these challenges, we first introduce a visual-native agent harness centered on an image bank reference protocol, which registers every tool-returned image as an addressable reference and makes intermediate visual evidence reusable by later tools. On top of this harness, On-policy Data Evolution (ODE) runs a closed-loop data generator that refines itself across rounds from rollouts of the policy being trained. This per-round refinement makes each round's data target what the current policy still needs to learn. The same framework supports both diverse supervised fine-tuning data and policy-aware reinforcement learning data curation, covering the full training lifecycle of the target agent. Across 8 multimodal deep search benchmarks, ODE improves the Qwen3-VL-8B agent from 24.9% to 39.0% on average, surpassing Gemini-2.5 Pro in standard agent-workflow setting (37.9%). At 30B, ODE raises the average score from 30.6% to 41.5%. Further analyses validate the effectiveness of image-bank reuse, especially on complex tasks requiring iterative visual refinement, while rollout-feedback evolution yields more grounded SFT traces and better policy-matched RL tasks than static synthesis.