发表机构
University of Pennsylvania; Shenzhen Campus of Sun Yat-sen University; Pace University(宾夕法尼亚大学; 中山大学深圳校区; 佩斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建可执行、可审计的合成Web环境,修复缺陷后用于训练Web智能体,提升了可行任务率与PPO策略性能,增强了向多个基准的迁移能力,为智能体学习提供了可靠训练基底。
AI 中文摘要
Web智能体有望实现复杂数字工作流程的自动化,但其训练仍受限于合成环境——这些环境看似合理,却隐藏着断开的链接、不一致的状态或不可行的任务。我们通过构建可执行、可审计且基于后端状态的合成Web环境,弥合了可扩展环境生成与可信智能体学习之间的差距。我们的框架将每个生成的网站表示为页面、导航链接、数据库记录、状态变更标记和任务约束的结构化支架,然后在策略训练前验证并修复结构、语义、一致性和可行性缺陷。交互过程中,普通UI转换以确定性方式执行,而持久的后端更新仅通过经验证的状态变更标记触发,从而能够基于经验证的任务进展谓词编译密集奖励。在涵盖六个领域的500个合成环境中,我们的方法减少了任务阻塞缺陷,将可行任务率从48.6%提升至94.8%,同时生成了更强的PPO策略,并在评估时无需调用LLM的情况下提升了向WebArena、WebShop和MiniWoB++的迁移性能。这些结果表明,经过验证的合成环境可作为紧凑Web智能体的可扩展且可靠的训练基底,将合成Web智能体学习从表面合理性转向可执行、基于状态的监督。
英文摘要
Web agents promise to automate complex digital workflows, but their training remains limited by synthetic environments that look plausible while hiding broken links, inconsistent states, or infeasible tasks. We address the gap between scalable environment generation and trustworthy agent learning by constructing synthetic web environments that are executable, auditable, and grounded in backend state. Our framework represents each generated website as a structured scaffold of pages, navigation links, database records, state-change markers, and task constraints, then verifies and repairs structural, semantic, consistency, and feasibility defects before policy training. During interaction, ordinary UI transitions are executed deterministically, while persistent backend updates are invoked only through validated state-change markers, enabling dense rewards compiled from verified task-progress predicates. Across 500 synthetic environments spanning six domains, our method reduces task-blocking defects and improves feasible-task rate from 48.6% to 94.8%, while producing stronger PPO policies and improving transfer to WebArena, WebShop, and MiniWoB++ without LLM calls at evaluation time. These results show that verified synthetic environments can serve as a scalable and reliable training substrate for compact web agents, shifting synthetic webagent learning from surface-level plausibility toward executable, state-grounded supervision.