渲染闭环:用于交互式网页开发的执行驱动智能体
Rendering-in-the-Loop: An Execution-Driven Agent for Interactive Web Development
- Baidu Inc.(百度公司)
- School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出执行驱动智能体RILA,通过AIV模块与ERS得分优化网页生成,构建执行验证数据流水线,在IWR-Bench上提升模型性能,使轻量Qwen3.5-9B超越Kimi-K2.6与GPT-5.5。
AI中文摘要:
多模态大语言模型在前端网页开发领域已取得显著进展,可从截图、交互视频等多模态参考生成交互式网页。但现有工作多侧重美学、布局相似度等视觉指标,却忽视了对交互功能这一更关键内容的验证。本文提出RILA,即一种将浏览器渲染纳入闭环的执行驱动智能体,可基于运行时交互反馈迭代编辑生成的代码。RILA引入了动作交互验证(AIV)模块,该模块会在生成的网页上复现参考交互轨迹,以收集基于执行的感知观测结果;还引入了执行感知渲染得分(ERS),该得分可联合衡量交互正确性与视觉保真度,用于指导迭代优化。我们进一步构建了经执行验证的数据合成流水线,该流水线可生成多样化的高质量训练数据,为推理时优化提供补充增益。在IWR-Bench基准测试中,RILA在各类基础模型上均实现了交互性与视觉保真度的同步提升。值得注意的是,借助我们的训练流水线,RILA将轻量型Qwen3.5-9B主干模型的性能从40.40%提升至57.52%,超过了大得多的单样本生成器,包括拥有1万亿参数的Kimi-K2.6(55.61%)以及专有模型GPT-5.5(55.74%)。
英文摘要:
Multimodal large language models have achieved remarkable progress in front-end web development, generating interactive webpages from multimodal references such as screenshots and interaction videos. However, existing work largely emphasizes visual metrics such as aesthetics and layout similarity, while overlooking the more critical validation of interactive functionality. We present RILA, an execution-driven agent that puts browser rendering in the loop, iteratively editing generated code from runtime interaction feedback. RILA introduces an Action Interaction Verification (AIV) module that replays the reference interaction trajectory on the generated webpage to collect grounded execution-aware observations, and an Execution-aware Rendering Score (ERS) that jointly measures interaction correctness and visual fidelity to guide iterative optimization. We further build an execution-verified data synthesis pipeline that produces diverse, high-quality training data, offering gains complementary to inference-time optimization. On IWR-Bench, RILA consistently improves both interaction and visual fidelity across foundation models. Notably, with our training pipeline, RILA lifts the compact Qwen3.5-9B backbone from 40.40% to 57.52%, surpassing far larger one-shot generators, including the 1T-parameter Kimi-K2.6 (55.61%) and the proprietary GPT-5.5 (55.74%).