arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

学习循环,而非仅页面:面向网页生成的基于执行的循环学习

Learning the Loop, Not Just the Page: Execution-Grounded Loop Learning for Web Generation

Yuxin Meng, Ruixu Zhang, Junjie Wang, Yuhan Suo, Yuhan Sun, Ruining Hu, Yiyao Yu, Yubin Wang, Shouwei Ruan, Bin Wang, Yue Liao, Yuxiang Zhang, Yujiu Yang

arXiv 2610.11543首次发表:更新:

发表机构

Tsinghua University; East China Normal University; Tongji University; Institute of Artificial Intelligence, Beihang University; National University of Singapore(清华大学; 华东师范大学; 同济大学; 北京航空航天大学人工智能研究院; 新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出基于执行的网页生成框架WebLoop,联合学习生成、评判与精炼,在WebRise、WebGen-Bench等基准上显著提升性能,且可迁移至大模型与多模态场景。

AI 中文摘要

功能性网页生成正越来越多地通过可执行奖励进行优化,但现有方法大多聚焦于最终页面的质量,而对不完善实现的诊断与修复过程探索不足。我们在此场景中识别出一个核心挑战:生成器(Generator)与精炼器(Refiner)生成带有直接环境奖励的可执行产物,而中间的评判器(Critic)会影响下游行为却无直接可执行结果。我们提出WebLoop,这是一个基于执行的框架,在共享策略中联合学习生成、评判与精炼。WebLoop训练了一个无执行的评判器,该评判器具备需求级可区分性与下游有用性的互补信号,先建立可靠诊断,再引入感知后果的信用分配,同时三者角色通过组相对策略学习进行联合优化。使用Qwen3.5-9B时,WebLoop在WebRise上达到41.5的整体得分,在WebGen-Bench上达到38.9%的准确率,分别比基础模型提升11.3和15.4个百分点。该提升可迁移至首次生成,在27B规模下仍保持,且能从纯文本训练泛化至多模态输入。控制分析进一步表明,该提升无法仅由额外的精炼步骤解释,凸显了学习评判器与循环本身的重要性。

英文摘要

Functional Web generation is increasingly optimized with executable rewards, yet existing methods largely focus on the quality of the final page and leave the process of diagnosing and repairing imperfect implementations underexplored. We identify a central challenge in this setting: the Generator and Refiner produce executable artifacts with direct environment rewards, whereas the intermediate Critic influences downstream behavior without a directly executable outcome. We introduce WebLoop, an execution-grounded framework that jointly learns generation, critique, and refinement within a shared policy. WebLoop trains an execution-free Critic with complementary signals for requirement-level discriminability and downstream helpfulness, first establishing reliable diagnosis and then introducing consequence-aware credit, while all three roles are jointly optimized with group-relative policy learning. With Qwen3.5-9B, WebLoop reaches 41.5 Overall on WebRise and 38.9% accuracy on WebGen-Bench, improving the base model by 11.3 and 15.4 points, respectively. The gains transfer to first-pass generation, persist at 27B scale, and generalize from text-only training to multimodal inputs. Controlled analyses further show that the improvement cannot be explained by an additional refinement pass alone, highlighting the importance of learning the Critic and the loop itself.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑