AI 中文总结
本研究提出自演化程序化评分器WebGrader,通过分离测试规划等环节生成准确奖励,训练8B策略在WebGen-Bench、WG-core-250数据集上的功能成功率及得分优于多款大语言模型。
AI 中文摘要
大语言模型越来越多地根据自然语言描述生成完整网站,强化学习已成为缩小其剩余功能差距的核心方法。该训练机制受限于奖励设计:手动编写的浏览器脚本可执行,但针对开放式需求编写成本高昂;而视觉语言模型(VLM)和图形用户界面(GUI)智能体评分器虽可扩展,但可能在观察到决定性状态前就给出判定结果。我们提出WebGrader,一种自演化程序化评分器,它能自主从每个网站请求中推导所需交互流,将每个流表示为可执行的流契约(Flow Contract),并将其执行结果作为强化学习(RL)的奖励。WebGrader在实时浏览器中实例化生成的项目,基于源代码和实时文档对象模型(DOM)确定目标动作,并沿同一浏览器轨迹收集视觉、DOM、响应和持久状态证据。随后,残差驱动的离线循环会发现可复用的验证器技能,在不相交的验证页面上对其进行筛选,并在策略训练前冻结已提升的技能图。通过将测试规划、动作定位、证据收集和语义判断分离,WebGrader仅在观察到请求的转换后才给出通过(Pass)判定。在WebGen-Bench上,WebGrader将一个80亿参数(8B)策略训练至52.01%的功能成功率,比匹配的外观加脚本奖励高出7.88个百分点,且优于o4-mini和DeepSeek-v4-flash;在WG-core-250上,该策略达到44.953的满分,且优于Qwen3-Coder-480B。
英文摘要
Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has become a central approach to closing their remaining functional gap. This training regime is bottlenecked by reward design. Hand-authored browser scripts are executable yet costly to write for open-ended requirements, while VLM and GUI-agent graders scale but may issue verdicts before observing the decisive state. We propose WebGrader, a self-evolving programmatic grader that autonomously derives the required interaction flows from each website request, represents each flow as an executable Flow Contract, and uses its execution outcome as an RL reward. WebGrader materializes the generated project in a live browser, grounds target actions against the source code and live DOM, and collects visual, DOM, response, and persistent-state evidence along the same browser trajectory. A residual-driven offline loop then discovers reusable verifier skills, screens them on disjoint validation pages, and freezes the promoted skill graph before policy training. By separating test planning, action grounding, evidence collection, and semantic judgment, WebGrader issues a Pass verdict only after observing the requested transition. On WebGen-Bench, WebGrader trains an 8B policy to a 52.01% functional success rate, outperforming a matched appearance-plus-script reward by 7.88 points and surpassing o4-mini and DeepSeek-v4-flash. On WG-core-250, the policy reaches a Full Score of 44.953 and surpasses Qwen3-Coder-480B.
Comments17 pages, 3 figures. Supplementary material is included in the main PDF