arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向Web智能体的判别式世界模型

Discriminative World Models for Web Agents

Kelvin Li, Dhruv Pendharkar, Anish Pahilajani, Chuyi Shang, Leon Oks, Leonid Karlinsky, Rogerio Feris, Trevor Darrell, Roei Herzig

arXiv 2609.02885首次发表:更新:

发表机构

University of California, Berkeley; MIT-IBM Watson AI Lab; Cal Poly San Luis Obispo; Xero(加州大学伯克利分校; 麻省理工学院-IBM沃森人工智能实验室; 加州理工州立大学圣路易斯奥比斯波分校; 赛罗公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对Web智能体世界模型训练与下游排序模型错位的问题,提出预测状态匹配训练目标,经多基准测试验证可提升动作排序及端到端任务成功率。

AI 中文摘要

现有Web智能体在测试阶段选择动作时,会利用世界模型对候选动作进行采样、预测对应的网页状态,再通过排序模型或过程奖励模型(Process Reward Model,PRM)对这些状态排序。这类世界模型通常通过监督式下一状态预测进行训练,以生成HTML或AXTree快照等固定表示。但该训练目标与下游排序模型存在错位,排序模型需要预测状态能区分不同候选动作对应的状态,才能准确打分。为解决这一问题,我们提出预测状态匹配(predicted-state matching)训练目标,要求预测表示能将真实结果状态与其他可选动作对应的状态区分开。我们使用源自WebArena Go-Browse轨迹的分支Web智能体数据集训练模型,该数据集的每个决策点都包含多个可选动作及其对应状态。在我们预留的预测状态匹配基准测试中,所提方法优于通过监督式下一状态预测训练的世界模型;在WebPRMBench上,与仅动作PRM及结合监督式下一状态世界模型的PRM相比,所提方法可提升PRM式动作排序效果;最后,在WebArena-Lite上,利用我们的世界模型进行测试阶段动作选择可提升端到端任务成功率。我们的项目页面可访问:this https URL。

英文摘要

Recent web agents use world models for test-time action selection by sampling candidate actions, predicting the resulting web states, and ranking them with a ranker model or a Process Reward Model (PRM). These world models are typically trained via supervised next-state prediction to generate fixed representations like HTML or AXTree snapshots. However, this objective is misaligned with the downstream ranker, which relies on predicted states being discriminative across candidates to accurately score them. To address this, we introduce predicted-state matching, a training objective where the predicted representation must distinguish the true resulting state from those reached by alternative actions. We train these models using a branching web-agent dataset derived from WebArena Go-Browse trajectories, where every decision point contains multiple alternative actions and their resulting states. Experiments on our held-out predicted-state matching benchmark show that our approach outperforms world models trained with supervised next-state prediction. We further show that our approach improves PRM-style action ranking on WebPRMBench compared with action-only PRMs and PRMs augmented with supervised-next-state world models. Finally, on WebArena-Lite, using our world model for test-time action selection improves end-to-end task success. Our project page is available at: https://dhruvpendharkar.github.io/dwm/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑