发表机构
MBZUAI; McGill University; CityUHK; UMass Boston(穆罕默德·本·扎耶德人工智能大学; 麦吉尔大学; 香港城市大学; 马萨诸塞大学波士顿分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SCAFFOLD框架,通过递归参数化技能抽象、MDL压缩和权重蒸馏,使视觉Web智能体自我改进,在多个基准上成功率提升11.1-17.2个百分点。
AI 中文摘要
Web智能体需要导航视觉丰富、长期跨度且在不同站点间变化的界面,然而以往大多数智能体仍孤立地学习每个任务,并丢弃它们积累的程序性知识。近期技能增强框架迈出了重要的第一步,但它们将技能库视为扁平或两层的提示侧缓存,且缺乏压缩冗余或递归组合技能的原则性机制。我们提出SCAFFOLD,一个用于视觉Web智能体的自我改进框架,它(i)在多实例抽象约束下,从成功轨迹中归纳参数化、可执行的技能;(ii)维护递归组合的层级结构,其中高层技能调用低层技能;(iii)通过最小描述长度(MDL)标准和行为等价性检查来压缩技能库;(iv)定期将技能增强轨迹蒸馏回模型权重,以内化这些抽象。在WebArena、VisualWebArena以及Online-Mind2Web的保留分割上,SCAFFOLD相比最强的技能增强基线,成功率提升了11.1至17.2个绝对百分点,并在五次自我改进迭代中展现出单调增益,且无技能库崩溃。我们在GitHub仓库中发布了代码和文档。
英文摘要
Web agents need to navigate visually rich, long-horizon interfaces that change across sites, yet most previous agents still learn each task in isolation and discard the procedural knowledge they accumulate. Recent skill-augmented frameworks take an important first step, but they treat the skill library as a flat or two-tier prompt-side cache and offer no principled mechanism for compressing redundancy or composing skills recursively. We introduce \textsc{Scaffold}, a self-improving framework for visual web agents that (i) induces parametric, executable skills from successful trajectories under a multi-instance abstraction constraint, (ii) maintains a recursively composed hierarchy in which higher-level skills invoke lower-level ones, (iii) compacts the library via a minimum-description-length (MDL) criterion and behavioral equivalence checking, and (iv) periodically distills skill-augmented trajectories back into model weights to internalize the abstractions. Across WebArena, VisualWebArena, and a held-out split of Online-Mind2Web, \textsc{Scaffold} improves success rate by $11.1$--$17.2$ absolute points over the strongest skill-augmented baseline and shows monotonic gains across five self-improvement iterations without library collapse. We release the code and documents in the Github \href{https://github.com/BokwaiHo/SCAFFOLD}{repository}.
CommentsAccepted by EMNLP 2026