arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28062cs.AI

WeAgent-MMSearch:面向多模态搜索智能体的原生文本-视觉交互

WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents

Zongkai Liu, Hui Zhang, Liqiang Niu, Zhen Cao, Han Li, Juntao Liu, Wenchao Chen, Chengduo Zhao, Chao Yu, Fandong Meng

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有多模态搜索智能体忽略图像、长程交互易失败的问题,提出 WeAgent-MMSearch 系统,通过 WeAgent-Harness 实现文本-视觉交互,经 FA-GSPO 后训练后,模型在多模态搜索基准上性能大幅提升。

中文摘要 AI 辅助

多模态搜索智能体利用开放网络中新兴及长尾证据扩展参数化知识,但现有多数智能体搜索环境仅将检索到的证据以文本形式呈现,且在后续上下文中忽略工具返回的图像,导致视觉基础轨迹沦为纯文本推理;长程交互还会加剧工具调用、响应长度、超时及预算失败问题,可能丢弃可挽救的轨迹、浪费 rollout 计算资源并干扰策略更新。为解决这些问题,我们引入 WeAgent-Harness,这是一种支持原生文本-视觉交互和运行时恢复的多模态智能体框架:检索到的图像会获得持久磁盘引用,使模型能在整个轨迹中检查、处理和引用这些图像。基于该框架,我们开发了 WeAgent-MMSearch,这是一个涵盖数据构建、智能体后训练及多模态 rollout 的集成系统。在数据构建阶段,强大的多模态大语言模型(MLLM)使用 WeAgent-Harness 发现、合成并验证 MMSearch 风格任务,同时收集专家轨迹;在后训练阶段,我们的故障感知 GSPO(FA-GSPO)可恢复可挽救的异常 rollout 并过滤无效 rollout,以改进有界多模态规划。此外,我们还推出了 VisTarget-Bench,这是一个包含 150 个任务、经人工验证的基准,每个任务的问题均配有预留目标图像,可区分图像检索失败与视觉感知失败。在 VisTarget-Bench 及七个公共基准上的评估显示,智能体后训练使平均分数提升了 19.22 个百分点,让我们的模型在性能上优于规模相近的开源模型,且可与参数数量约为其十倍的模型相媲美。

英文摘要

Multimodal search agents extend parametric knowledge with newly emerging and long-tail evidence from the open web. Yet many existing agentic search environments often expose retrieved evidence only as text and omit tool-returned images from subsequent context, reducing visually grounded trajectories to text-only reasoning. Long-horizon interaction also compounds tool-call, response-length, timeout, and budget failures, which can discard salvageable trajectories, waste rollout computation, and disturb policy updates. To address these issues, we introduce WeAgent-Harness, a multimodal agentic harness that supports native text-vision interaction and runtime recovery. Retrieved images receive persistent disk references, allowing the model to inspect, process, and cite them throughout the trajectory. Based on this harness, we develop WeAgent-MMSearch, an integrated system spanning data construction, agentic post-training, and multimodal rollout. For data construction, a strong MLLM uses WeAgent-Harness to discover, synthesize, and verify MMSearch-style tasks and collect expert trajectories. During post-training, our Failure-Aware GSPO (FA-GSPO) recovers salvageable abnormal rollouts and filters invalid ones to improve bounded multimodal planning and search. We also introduce VisTarget-Bench, a 150-task human-verified benchmark that pairs each question with a held-out target image, distinguishing image-retrieval failures from visual-perception failures. Evaluation on VisTarget-Bench and seven public benchmarks shows that agentic post-training improves the average score by 19.22 points, enabling our model to outperform similarly sized open-source models and rival models with roughly ten times its parameter count.

发表机构

  • Weixin AI, Tencent(腾讯微信人工智能)
  • Sun Yat-sen University(中山大学)

机构由 AI 辅助整理,请以论文原文为准。

↑