arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MAG:用于多模态动作与引导生成的网络智能体基准测试与工具包

MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation

Chengguang Gan, Hanjun Wei, Yunhao Liang, Zhixi Cai, Qinghao Zhang, Shiwen Ni

arXiv 2607.10079首次发表:更新:

发表机构

University of Chinese Academy of Sciences; Monash University; Pusan National University; Shenzhen University of Advanced Technology(中国科学院大学; 莫纳什大学; 釜山国立大学; 深圳先进技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

介绍MAG这一网络智能体基准测试,统一任务执行与引导写作,有基于截图的定位方案和完整工具包。用其评估模型并详细分析,还设计GRPO训练方法,提升智能体成功率与引导质量,指出当前模型任务完成率低,为后续研究提供方向。

AI 中文摘要

数字采用平台(DAPs)广泛用于网络系统,帮助用户在页面内操作,但完成实际任务需跨页面状态执行一系列动作。此前研究将网络智能体动作与引导文本生成视为两个独立问题,且多使用文本页面表示训练模型。本文介绍MAG,首个将任务执行与引导写作统一为多模态动作与引导任务的基准测试,有两种基于截图的定位方案。还构建了完整工具包,涵盖借助大语言模型辅助注释和人工验证、训练、实时环境评估及动作与引导联合指标。用此工具包评估前沿API模型和开放多模态模型并详细分析。最后设计了GRPO训练方法,使监督式9B智能体成功率从6.9%提高到13.2%,同时提升引导质量。即便最强模型也只能完成不到40%的任务,为未来研究留足空间。

英文摘要

Digital Adoption Platforms (DAPs) are embedded overlays widely used on web systems to guide users through operations inside a page, helping them get started with unfamiliar interfaces quickly. Completing a real task, however, rarely means clicking a few buttons on a single page: it takes a sequence of actions that unfolds across changing page states. Prior studies have also treated automated web agent actions and guide text generation as two separate problems, and most of them feed models textual page representations such as the DOM or accessibility trees rather than the rendered screens that humans actually operate on. In this work we introduce MAG, the first benchmark that unifies task execution and guide writing into a single Multimodal Action and Guide task, with two grounding schemes over screenshots: Set-of-Mark element selection and raw pixel coordinates. We further build a complete harness for this compound task, covering annotation with LLM assistance and human verification, training, evaluation in live environments, and joint metrics for actions and guides. With this harness we evaluate frontier API models and open multimodal models, and report detailed analyses. Finally, we design a GRPO training method augmented with expert trajectories, which nearly doubles the success rate of a supervised 9B agent (from 6.9% to 13.2%) and improves guide quality at the same time. Even the strongest model completes fewer than 40% of the tasks, leaving ample room for future research.

Comments8 pages main text, 21 pages total including appendices; 11 figures, 7 tables, 2 algorithms. Benchmark, harness, and model checkpoints to be released

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑