引导引发行动:多模态网页智能体中引导-行动相互增强的离线研究
Guides That Cause Actions: An Offline Study of Guide-Action Mutual Reinforcement in Multimodal Web Agents
浏览论文内容
中文总结 AI 辅助
提出离线基准WebMRE,研究网页智能体中引导句与行动的相互增强效应,证明引导是因果通道,微调模型在离线指标上超越多个前沿模型。
中文摘要 AI 辅助
网页智能体通常在线下环境中进行评估,其中环境状态和评判模型在不同运行之间漂移,因此同一检查点很少能重现相同的分数,这使得对训练现象进行受控研究变得不切实际。我们提出了WebMRE,一个包含541个任务和5,293个步骤的离线基准,这些任务和步骤源自成功的WebArena轨迹,具有完全审计的测试标签和确定性协议,该协议在每次运行中无需任何环境即可对检查点进行相同评分。每个步骤将面向人类的引导句与有根据的行动配对,从而首次研究网页智能体中两者之间的相互增强效应。在三个种子上平均,该效应在两种解码顺序下对两个模型均成立,并随规模增长:联合解码引导句将元素选择相对于仅行动参考提高了Qwen3.5-4B的0.9和0.2个百分点,以及Qwen3.5-9B的1.7和2.2个百分点。中介分析表明,引导是因果通道而非评论:强制将黄金引导作为解码前缀将行动准确率从.422提高到.684,另一个步骤的引导将其降至.055,而重命名目标的释义仍能恢复一半的增益,因此该通道承载指令含义而不仅仅是标签字符串。同一通道产生离线奖励,只有可重放的协议才能计算,尽管从强检查点优化它尚未带来增益。我们微调的模型在零样本运行下,在每项离线指标上均优于GPT-5.5、Claude Opus 4.8和Gemini 3.5 Flash。
英文摘要
Web agents are usually evaluated in live environments, where environment state and judge models drift between runs, so the same checkpoint rarely reproduces the same score, making controlled studies of training phenomena impractical. We present WebMRE, an offline benchmark of 541 tasks and 5,293 steps derived from successful WebArena trajectories, with fully audited test labels and a deterministic protocol that scores a checkpoint identically on every run without any environment. Each step pairs a human oriented guide sentence with a grounded action, enabling the first study of the mutual reinforcement effect between them in web agents. Averaged over three seeds the effect holds for both models in both decoding orders and grows with scale: jointly decoding a guide lifts element selection over an action only reference by 0.9 and 0.2 points for Qwen3.5-4B and by 1.7 and 2.2 points for Qwen3.5-9B. A mediation analysis shows that the guide is a causal channel rather than commentary: forcing the gold guide as a decoding prefix lifts action accuracy from .422 to .684, another step's guide collapses it to .055, and a paraphrase that renames the target still recovers half of the gain, so the channel carries instruction meaning and not only the label string. The same channel yields an offline reward that only a replayable protocol makes computable, though optimizing it from a strong checkpoint brings no gain yet. Our fine tuned models outperform GPT-5.5, Claude Opus 4.8, and Gemini 3.5 Flash, run zero shot, on every offline metric.
发表机构
- University of Chinese Academy of Sciences(中国科学院大学)
- Pusan National University(釜山国立大学)
- Shenzhen University of Advanced Technology(深圳理工大学)
机构由 AI 辅助整理,请以论文原文为准。