arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UNIBROWSE:用于多模态浏览比较的数据到智能体框架

UNIBROWSE: A Data-to-Agent Framework for Multimodal BrowseComp

Xiyu Wei, Qingwei Zong, Zhuocheng Yu, Sujian Li

arXiv 2607.10557首次发表:更新:

发表机构

Key Laboratory of Computational Linguistics, MOE, Peking University; School of Software and Microelectronics, Peking University; School of Computer Science, Peking University(教育部计算语言学重点实验室,北京大学; 北京大学软件与微电子学院; 北京大学计算机科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多模态浏览比较任务,提出UNIBROWSE统一数据管道,同时生成三种模式训练数据,增强知识图谱,引入探索度度量。经此训练的35B规模智能体在多模态浏览比较基准测试中性能领先,超多个闭源智能体工作流程。

AI 中文摘要

多模态浏览比较任务要求智能体结合感知、工具使用以及对动态网页内容的长期推理,这对其处理组合结构、开放世界不确定性和跨扩展交互的多模态集成能力提出挑战。现实世界的多模态浏览涉及三种不同的信息流模式,而现有数据构建方法仅涵盖了其中两种,导致文本到图像模式未得到充分解决,限制了智能体的通用性和鲁棒性。我们引入了UNIBROWSE,这是一个统一的数据管道,首次同时生成涵盖所有三种模式的训练数据,通过实时网络检索增强精选知识图谱以提高保真度,并引入一种新的探索度度量来过滤低信号实例以进行高效强化学习。通过这个管道,我们生成了高质量的冷启动工具使用轨迹和富含探索的问答对,并通过监督微调训练了一个35B规模的智能体。由此产生的UNIBROWSE智能体在多模态浏览比较基准测试中取得了领先性能,在五个不同基准测试中的平均准确率达到54.4,比其基础模型Qwen3.5 - 35B - A3B提高了10.5个百分点,超过了多个闭源智能体工作流程。

英文摘要

Multimodal BrowseComp tasks require agents to combine perception, tool use, and long-horizon reasoning over dynamic web content, challenging their ability to handle compositional structure, open-world uncertainty, and multimodal integration across extended interactions. Crucially, real-world multimodal browsing involves three distinct information-flow patterns: text-only, image-to-text, and text-to-image, yet existing data construction methods cover only the text-only and image-to-text patterns, leaving text-to-image largely unaddressed and limiting agent generality and robustness. We introduce UNIBROWSE, a unified data pipeline that for the first time simultaneously generates training data covering all three patterns, augments curated knowledge graphs with live web retrieval for improved fidelity, and introduces a novel metric of exploration degree to filter low-signal instances for efficient reinforcement learning. Through this pipeline, we produce high-quality cold-start tool-use trajectories and exploration-rich QA pairs, and train a 35B-scale agent via supervised fine-tuning and exploration-aware RL.The resulting UNIBROWSE agent achieves state-of-the-art performance on multimodal BrowseComp benchmarks, attaining an average accuracy of 54.4 across five diverse benchmarks -- an improvement of 10.5 points over its base model Qwen3.5-35B-A3B -- and surpassing serveral closed-source agent workflows such as GPT-5 (42.9), Gemini-2.5 Pro (44.8), and Gemini-2.5 Flash (41.3).

Comments17 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑