arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00335cs.AIcs.CLcs.LG

RMSWeb:面向Web智能体强化学习的反思、失败模式挖掘与Salvage-DS

RMSWeb: Reflection, Failure-Mode Mining, and Salvage-DS for Web Agent Reinforcement Learning

Chengbo Liu, Lifang Zhou, Ruijie Yan, Pei Tan, Ao Sun, Haojun Huang, Guichun Hua, Sining Wei, Yining Chen, Yingying He, Yutao Xie

首次发表
浏览论文内容

中文总结 AI 辅助

RMSWeb是针对Qwen3-VL-Instruct 8B和32B模型的Web智能体强化学习方案,通过三部分技术提升训练效率与性能,在多个Web基准数据集上较SFT取得显著提升,且8B模型获同规模开源模型最优Online-Mind2Web结果。

中文摘要 AI 辅助

紧凑的Web智能体可降低部署成本,但训练它们在数据收集和监督微调(SFT)后的强化学习(RL)阶段均面临挑战:成功轨迹收集成本高昂且常包含低效绕行;SFT后全轨迹语料以常规状态为主;将组相对RL应用于Web动作时,设计不当的动作级奖励会产生弱或误导性的相对更新,被判定不适合此类更新的组则无法获得回退学习信号。本文提出RMSWeb,这是针对Qwen3-VL-Instruct 8B和32B模型的三部分方案:反思条件重试可提升收集效率并缩短成功轨迹;失败模式挖掘将离线RL聚焦于SFT策略暴露的关键状态;Salvage-DS结合动作语义极化奖励、对比-能力门控动态采样,以及针对被拒组的仅动作锚点。使用反思收集数据训练的策略,在已解决任务上的动作步骤最多减少19.7%。在WebVoyager、Online-Mind2Web和WebTailBench数据集上,RMSWeb相比SFT在8B模型上提升2.4-7.0个百分点,在32B模型上提升1.2-7.7个百分点。本文的8B模型在对比的同规模开源权重模型中,取得了Online-Mind2Web报告的最强结果,且在WebVoyager和WebTailBench上实现了领先的准确率-成本权衡,但需注意外部评估协议存在差异。

英文摘要

Compact web agents can reduce deployment cost, but training them poses challenges in both data collection and post-SFT reinforcement learning (RL). Successful trajectories are expensive to collect and often contain inefficient detours. After supervised fine-tuning (SFT), full trajectory corpora are dominated by routine states; moreover, when group-relative RL is applied to web actions, inadequately designed action-level rewards can yield weak or misleading relative updates, while groups rejected as unsuitable for such updates receive no fallback learning signal. We present RMSWeb, a three-part recipe for Qwen3-VL-Instruct at 8B and 32B. Reflection-conditioned retries increase collection yield and shorten successful trajectories; failure-mode mining concentrates offline RL on critical states exposed by the SFT policy; and Salvage-DS combines an action-semantic polarized reward, contrast-and-competence-gated dynamic sampling, and an action-only anchor for rejected groups. Policies trained with reflection-collected data use up to 19.7% fewer action steps on solved tasks. On WebVoyager, Online-Mind2Web, and WebTailBench, RMSWeb improves over SFT by 2.4-7.0 points at 8B and 1.2-7.7 points at 32B. Our 8B model also achieves the strongest reported Online-Mind2Web result among similarly sized open-weight models in our comparison and a leading reported accuracy-cost trade-off on WebVoyager and WebTailBench, with the caveat that external evaluation protocols differ.

补充信息

↑