arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32351cs.AI

从轨迹到基于环境的偏好:通过交互元素图进行Web PRM的过程偏好合成

From Trajectories to Grounded Preferences: Process Preference Synthesis via Interaction Element Graphs for Web PRMs

Yangzhe Peng, Xiaoyang Wang, Yiyang Zhao, Lijun Wu, Kun He

首次发表
浏览论文内容

中文总结 AI 辅助

针对Web PRM偏好数据中GMCP稀缺问题,提出SURFPRM框架,通过交互元素图合成对比负动作,将GMCP比例从24.19%提升至74.60%,显著提升PRM性能及下游任务成功率。

中文摘要 AI 辅助

比较过程奖励模型(PRMs)通过评估候选动作之间基于状态的偏好,为自主Web智能体提供关键的步骤级指导。然而,现有的通过多策略采样合成的偏好训练数据严重缺乏基于环境的最小对比对(GMCPs)——其中竞争候选动作针对真实的页面元素且具有相同的动作类型。在代表性的基线偏好数据(即WebArbiter)中,GMCPs仅占24.19%,导致PRMs在训练期间偏向依赖浅层捷径(如元素幻觉和动作类型不匹配),而非获取真正的上下文决策语义。为解决这些挑战,我们提出了SURFPRM,一个用于比较Web PRMs的图引导的过程偏好合成框架。SURFPRM将Web演示结构化为一个持久的交互元素图,该图充当基于环境的负动作提议机制,系统地跨空间、时间和时空混淆轴合成对比性负动作。这将GMCP比例从24.19%提升至74.60%,生成了精选的SURFPRM-DATA数据集。在六个开源骨干模型(3B至9B参数)上,基于SURFPRM-DATA训练的PRMs在WEBPRMBENCH上平均优于基线训练的模型,并与领先的专有LLMs相媲美。在下游的奖励引导轨迹搜索中,在WEBARENA-LITE上,SURFPRM为GPT-4o(+14.21%)和GPT-4o-mini(+12.83%)策略提供步骤级指导,在复杂Web任务成功率上带来了显著提升。

英文摘要

Comparative Process Reward Models (PRMs) provide critical step-level guidance for autonomous web agents by evaluating state-conditioned preferences between candidate actions. However, existing preference training data synthesized via multi-policy sampling suffers from a severe scarcity of Grounded Minimal Contrastive Pairs (GMCPs)-where competing candidates target genuine on-page elements with identical action types. In representative baselines preference data (namely, WebArbiter), GMCPs account for merely 24.19%, biasing PRMs during training to rely on shallow shortcuts (such as element hallucinations and action type mismatches) rather than acquiring genuine contextual decision semantics. To address these challenges, we propose SURFPRM, a graph-guided process preference synthesis framework for comparative Web PRMs. SURFPRM structures web demonstrations into a persistent Interaction Element Graph that acts as an environment-grounded negative action proposal mechanism, systematically synthesizing contrastive negative actions across spatial, temporal, and spatiotemporal confusion axes. This elevates the GMCP proportion from 24.19% to 74.60%, producing the curated SURFPRM-DATA dataset. Across six open-source backbones (3B to 9B parameters), PRMs trained on SURFPRM-DATA outperform baseline-trained models on average on WEBPRMBENCH and rival leading proprietary LLMs. In downstream reward-guided trajectory search on WEBARENA-LITE, SURFPRM provides step-level guidance for both GPT-4o (+14.21%) and GPT-4o-mini (+12.83%) policies, yielding substantial improvements in complex web task success rates.

↑