arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30530cs.CLcs.SE

WebWorld:将浏览器作为自改进网页代码的世界模型

WebWorld: The Browser as a World Model for Self-Improving Web Code

Jiajun Wu, Jian Yang, Yaxin Du, Wei Zhang, Haowen Wang, Junhang Cheng, Yuxuan Zhang, Tuney Zheng, Xianglong Liu, Ming Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出WebWorld,将浏览器作为VLM无法欺骗的网页代码世界模型,通过浏览器签发合格证书的机制实现VLM自主交互与监督,使WebWorld-27B在两个基准任务上实现显著性能提升,达到强前沿系统水平。

中文摘要 AI 辅助

基于视觉语言模型(VLM)的网页代码自改进存在结构性缺陷:提出修复方案的模型同时也是评判该方案的模型,而在该评判标准下的视觉合理性,与网页实际是否可用的关联度极低。该循环缺少一个VLM无法欺骗的对象,而浏览器正是这样的对象:它是HTML制品在用户操作下行为的确定性、可执行模拟器,本质上就是网页代码的世界模型。我们提出WebWorld,该接口让预训练VLM能自主与这个作为世界模型的浏览器交互,并决定哪些交互可作为监督信号。每一轮中,VLM生成一份批评意见,规划器将其编译为带类型的交互契约;浏览器会重新执行候选代码,仅当满足目标进展且保留所有先前验证过的功能时,才会签发合格证书;经认证的转换会累积为质量棘轮,这也是SFT(监督微调)导出时唯一可见的内容。在匹配训练下,WebWorld-27B在HTMLBench-400上较Raw-27B提升5.3个百分点,在MiniAppBench-Val上提升14.9个百分点,在交互式HTML生成上达到Kimi-K2.6、GPT-5.4等强前沿系统的水平。同等规模的消融实验表明,浏览器背书的准入机制是性能提升的关键:若无合格证书,匹配训练的9B模型的提升几乎消失。

英文摘要

VLM-driven self-improvement of web code has a structural flaw: the model that proposes the repair is the model that judges it, and visual plausibility under that judge is a poor proxy for whether the page actually works. What the loop is missing is a counterparty the VLM cannot fool, and the browser already is that counterparty: a deterministic, executable simulator of how an HTML artifact behaves under user actions, and in everything but name a world model for web code. We present WebWorld, the interface that lets a VLM prior interact with this browser-as-world-model autonomously and decides which interactions become supervision. Each round, the VLM emits a critique that the planner compiles into a typed interaction contract; the browser re-executes the candidate and issues an acceptance certificate only when both target progress and preservation of every previously verified capability hold; certified transitions accumulate as a quality ratchet that is the only thing the SFT export ever sees. Under matched training, WebWorld-27B improves Raw-27B by 5.3 points on HTMLBench-400 and 14.9 points on MiniAppBench-Val, and reaches the level of strong frontier systems such as Kimi-K2.6 and GPT-5.4 on interactive HTML generation. Equal-size ablations show that browser-backed admission carries the gain: without the certificate, the matched 9B lift nearly disappears.

发表机构

  • Beihang University(北京航空航天大学)
  • Shanghai Jiao Tong University(上海交通大学)
  • IQuest Research(IQuest研究院)
  • Langboat

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑