arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24848cs.CL

BrowserForge:通过并行浏览器沙盒扩展网页交互序列数据规模

BrowserForge: Scaling Web Episode via Parallel Browser Sandboxes

Fei Tang, Huawen Shen, Zhiqiong Lu, Zhengxi Lu, Pengyuan Lyu, Chengquan Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen

首次发表
浏览论文内容

中文总结 AI 辅助

提出BrowserForge框架,通过并行浏览器沙盒从开放网络生成20余万条不同网站的网页交互轨迹,微调多模态模型后在两个基准任务上性能显著提升。

中文摘要 AI 辅助

基于渲染像素行动的网页智能体,避免了读取页面HTML或可访问性树的脆弱性及高额token成本,但训练这类智能体依赖大量高质量交互轨迹,如何规模化生成此类数据仍是未解决的问题。公开数据集通常仅包含来自固定且狭窄网站集合的数千条轨迹,即便是近期的自动化合成流水线也受限于预定义的网站列表或教程源,导致智能体接触到的不同网站数量几乎没有增长。我们提出BrowserForge,一个通过在开放网络上并行驱动多个浏览器沙盒来规模化生成网页交互数据的框架。BrowserForge包含三个组件:开放网络源阶段,使智能体接触数十万真实、可公开访问的网站;沙盒集群管理器,以高利用率调度数百个并发浏览器;以及一个提议者-求解者双智能体循环,将原始页面转化为可执行任务,随后为其收集经过验证的轨迹。一个规则加模型的清洗流水线会移除失败的运行,并将幸存的推理改写为单一的统一思维链风格。页面结构(如可访问性树)仅用作合成时的信号;我们训练并发布的智能体仅从截图行动。生成的语料库包含203238条轨迹,每条来自不同网站,比现有轨迹数据集更大、更多样化。在该语料库上对紧凑多模态模型进行微调,使其在实时Online-Mind2Web上的成功率从25.66%提升至33.33%,并持续提升静态Multimodal-Mind2Web上的步骤准确率,且增益随语料库规模扩大而增长。受控分析进一步证实,开放网络源和广泛的网站覆盖是观察到的性能提升的关键因素。

英文摘要

Web agents that act from rendered pixels avoid the fragility and heavy token cost of reading a page's HTML or accessibility tree, but training them depends on large amounts of high-quality interaction trajectories, and how to produce such data at scale remains an open problem. Public datasets typically contain only a few thousand trajectories drawn from a fixed and narrow set of websites, and even recent automated synthesis pipelines stay bound to predefined site lists or tutorial sources, so the number of distinct websites the agent ever sees barely grows. We present BrowserForge, a framework that generates web interaction data at scale by driving many browser sandboxes in parallel over the open web. BrowserForge couples three components: an open-web sourcing stage that exposes the agent to hundreds of thousands of real, openly reachable websites; a sandbox cluster manager that schedules hundreds of concurrent browsers with high utilization; and a Proposer-Solver dual-agent loop that turns a raw page into an executable task and then collects a verified trajectory for it. A rule-plus-model cleaning pipeline removes failed runs and rewrites the surviving reasoning into a single unified chain-of-thought style. Page structure such as the accessibility tree is used only as a synthesis-time signal; the agent we train and release acts purely from the screenshot. The resulting corpus contains 203,238 trajectories, each collected from a distinct website, larger and more diverse than prior trajectory datasets. Fine-tuning a compact multimodal model on this corpus raises its success rate on the live Online-Mind2Web from 25.66% to 33.33% and consistently improves step accuracy on the static Multimodal-Mind2Web, with the gain growing as the corpus scales. Controlled analyses further confirm that open-web sourcing and broad website coverage are key contributors to the observed improvement.

发表机构

  • Zhejiang University(浙江大学)
  • Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

↑