arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03474cs.CV

MT-Web2Code:面向多轮区域重构与局部修改的编码智能体基准测试

MT-Web2Code: Benchmarking Coding Agents on Multi-Turn Regional Reconstruction and Localized Modification

Qiming Li, Shujie Hu, Haohan Liu, Xiaocheng Feng, Songxiang Liu, Guanglu Wan

首次发表
浏览论文内容

中文总结 AI 辅助

MT-Web2Code是首个面向多轮区域重构与局部修改的多模态编码基准,实验发现现有编码智能体存在多轮错误滚雪球等问题,其评估指标或助力相关研究。

中文摘要 AI 辅助

大型视觉语言模型(LVLMs)在网页UI生成方面已展现出令人印象深刻的能力。然而,现有基准大多聚焦于从零开始的单轮全页面生成,忽略了实际前端工程的迭代工作流程——开发者会在现有代码库中反复重构缺失区域并修改局部元素。为填补这一空白,我们推出了MT-Web2Code,这是首个针对多轮宏观层面区域重构与微观层面局部修改的多模态编码基准,包含覆盖16个垂直领域的102个任务。为构建无需高成本轮级人工标注的确定性修复轨迹,我们开发了可扩展的反向损坏轨迹引擎,该引擎会向黄金页面迭代注入结构和样式缺陷。我们进一步提出了双轴评估协议,用于衡量目标区域的保真度与未受影响内容的保留情况,其中区域重构通过基于VLM的5维度准则评估,局部修改则通过确定性像素级对齐评估。对13个前沿编码智能体的实验显示,当前智能体在忠实重构目标区域的同时保留未受影响内容方面存在困难,在局部编辑上缺乏细粒度的视觉-代码对齐,且在多轮过程中存在错误滚雪球问题。除基准测试外,我们的确定性评估指标提供了细粒度反馈信号,可能会推动未来迭代式UI编码智能体的研究。我们的评估代码和数据将很快发布。

英文摘要

Recent advances in Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in web UI generation. However, existing benchmarks predominantly focus on single-turn full-page generation from scratch, overlooking the iterative workflow of real-world frontend engineering, where developers repeatedly reconstruct missing regions and modify localized elements within existing codebases. To bridge this gap, we introduce MT-Web2Code, the first multimodal coding benchmark for multi-turn Macro-Level Regional Reconstruction and Micro-Level Localized Modification, which contains 102 tasks spanning 16 vertical domains. To construct deterministic repair trajectories without costly turn-level human annotation, we develop a scalable Reverse-Corruption Trajectory Engine that iteratively injects structural and stylistic defects into golden pages. We further propose a dual-axis evaluation protocol that measures target-region fidelity and the preservation of unaffected content, where regional reconstruction is assessed by a 5-dimensional VLM-based rubric and localized modification by deterministic pixel-grounded alignment. Experiments on 13 frontier coding agents reveal that current agents struggle to faithfully reconstruct target regions while preserving unaffected content, lack fine-grained visual-code alignment for localized edits, and suffer from error snowballing over multiple turns. Beyond benchmarking, our deterministic evaluation metrics provide fine-grained feedback signals that may facilitate future research on training iterative UI coding agents. Our evaluation code and data will soon be released.

↑