RMSWeb:面向Web智能体强化学习的反思、失败模式挖掘与Salvage-DS
RMSWeb: Reflection, Failure-Mode Mining, and Salvage-DS for Web Agent Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
RMSWeb是针对Qwen3-VL-Instruct 8B和32B模型的Web智能体强化学习方案,通过三部分技术提升训练效率与性能,在多个Web基准数据集上较SFT取得显著提升,且8B模型获同规模开源模型最优Online-Mind2Web结果。
中文摘要 AI 辅助
紧凑的Web智能体可降低部署成本,但训练它们在数据收集和监督微调(SFT)后的强化学习(RL)阶段均面临挑战:成功轨迹收集成本高昂且常包含低效绕行;SFT后全轨迹语料以常规状态为主;将组相对RL应用于Web动作时,设计不当的动作级奖励会产生弱或误导性的相对更新,被判定不适合此类更新的组则无法获得回退学习信号。本文提出RMSWeb,这是针对Qwen3-VL-Instruct 8B和32B模型的三部分方案:反思条件重试可提升收集效率并缩短成功轨迹;失败模式挖掘将离线RL聚焦于SFT策略暴露的关键状态;Salvage-DS结合动作语义极化奖励、对比-能力门控动态采样,以及针对被拒组的仅动作锚点。使用反思收集数据训练的策略,在已解决任务上的动作步骤最多减少19.7%。在WebVoyager、Online-Mind2Web和WebTailBench数据集上,RMSWeb相比SFT在8B模型上提升2.4-7.0个百分点,在32B模型上提升1.2-7.7个百分点。本文的8B模型在对比的同规模开源权重模型中,取得了Online-Mind2Web报告的最强结果,且在WebVoyager和WebTailBench上实现了领先的准确率-成本权衡,但需注意外部评估协议存在差异。
英文摘要
Compact web agents can reduce deployment cost, but training them poses challenges in both data collection and post-SFT reinforcement learning (RL). Successful trajectories are expensive to collect and often contain inefficient detours. After supervised fine-tuning (SFT), full trajectory corpora are dominated by routine states; moreover, when group-relative RL is applied to web actions, inadequately designed action-level rewards can yield weak or misleading relative updates, while groups rejected as unsuitable for such updates receive no fallback learning signal. We present RMSWeb, a three-part recipe for Qwen3-VL-Instruct at 8B and 32B. Reflection-conditioned retries increase collection yield and shorten successful trajectories; failure-mode mining concentrates offline RL on critical states exposed by the SFT policy; and Salvage-DS combines an action-semantic polarized reward, contrast-and-competence-gated dynamic sampling, and an action-only anchor for rejected groups. Policies trained with reflection-collected data use up to 19.7% fewer action steps on solved tasks. On WebVoyager, Online-Mind2Web, and WebTailBench, RMSWeb improves over SFT by 2.4-7.0 points at 8B and 1.2-7.7 points at 32B. Our 8B model also achieves the strongest reported Online-Mind2Web result among similarly sized open-weight models in our comparison and a leading reported accuracy-cost trade-off on WebVoyager and WebTailBench, with the caveat that external evaluation protocols differ.