WeAgent-MMSearch:面向多模态搜索智能体的原生文本-视觉交互
WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents
浏览论文内容
中文总结 AI 辅助
针对现有多模态搜索智能体忽略图像、长程交互易失败的问题,提出 WeAgent-MMSearch 系统,通过 WeAgent-Harness 实现文本-视觉交互,经 FA-GSPO 后训练后,模型在多模态搜索基准上性能大幅提升。
中文摘要 AI 辅助
多模态搜索智能体利用开放网络中新兴及长尾证据扩展参数化知识,但现有多数智能体搜索环境仅将检索到的证据以文本形式呈现,且在后续上下文中忽略工具返回的图像,导致视觉基础轨迹沦为纯文本推理;长程交互还会加剧工具调用、响应长度、超时及预算失败问题,可能丢弃可挽救的轨迹、浪费 rollout 计算资源并干扰策略更新。为解决这些问题,我们引入 WeAgent-Harness,这是一种支持原生文本-视觉交互和运行时恢复的多模态智能体框架:检索到的图像会获得持久磁盘引用,使模型能在整个轨迹中检查、处理和引用这些图像。基于该框架,我们开发了 WeAgent-MMSearch,这是一个涵盖数据构建、智能体后训练及多模态 rollout 的集成系统。在数据构建阶段,强大的多模态大语言模型(MLLM)使用 WeAgent-Harness 发现、合成并验证 MMSearch 风格任务,同时收集专家轨迹;在后训练阶段,我们的故障感知 GSPO(FA-GSPO)可恢复可挽救的异常 rollout 并过滤无效 rollout,以改进有界多模态规划。此外,我们还推出了 VisTarget-Bench,这是一个包含 150 个任务、经人工验证的基准,每个任务的问题均配有预留目标图像,可区分图像检索失败与视觉感知失败。在 VisTarget-Bench 及七个公共基准上的评估显示,智能体后训练使平均分数提升了 19.22 个百分点,让我们的模型在性能上优于规模相近的开源模型,且可与参数数量约为其十倍的模型相媲美。
英文摘要
Multimodal search agents extend parametric knowledge with newly emerging and long-tail evidence from the open web. Yet many existing agentic search environments often expose retrieved evidence only as text and omit tool-returned images from subsequent context, reducing visually grounded trajectories to text-only reasoning. Long-horizon interaction also compounds tool-call, response-length, timeout, and budget failures, which can discard salvageable trajectories, waste rollout computation, and disturb policy updates. To address these issues, we introduce WeAgent-Harness, a multimodal agentic harness that supports native text-vision interaction and runtime recovery. Retrieved images receive persistent disk references, allowing the model to inspect, process, and cite them throughout the trajectory. Based on this harness, we develop WeAgent-MMSearch, an integrated system spanning data construction, agentic post-training, and multimodal rollout. For data construction, a strong MLLM uses WeAgent-Harness to discover, synthesize, and verify MMSearch-style tasks and collect expert trajectories. During post-training, our Failure-Aware GSPO (FA-GSPO) recovers salvageable abnormal rollouts and filters invalid ones to improve bounded multimodal planning and search. We also introduce VisTarget-Bench, a 150-task human-verified benchmark that pairs each question with a held-out target image, distinguishing image-retrieval failures from visual-perception failures. Evaluation on VisTarget-Bench and seven public benchmarks shows that agentic post-training improves the average score by 19.22 points, enabling our model to outperform similarly sized open-source models and rival models with roughly ten times its parameter count.
发表机构
- Weixin AI, Tencent(腾讯微信人工智能)
- Sun Yat-sen University(中山大学)
机构由 AI 辅助整理,请以论文原文为准。