arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12518cs.SE

它是否在所有环境中都能渲染?对多模态大语言模型(MLLM)生成网页的跨环境兼容性研究

Does It Render Everywhere? A Study of Cross-Environment Compatibility in MLLM-Generated Webpages

Ziyun Guo, Jingyu Xiao, Yuqiang Sun, Yintong Huo

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对MLLM生成网页的跨环境兼容性问题,构建WebCompat数据集并开发XCompat检测器,发现68%的网页存在兼容性问题,XCompat的F1分数达0.903且性能优于现有工具与基线方法。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)已越来越多地被用于从视觉设计(如截图)自动生成网页。然而,现有评估仅局限于在固定的浏览器-设备配置下进行视觉保真度评估,这种设置忽略了面向实际部署的跨环境渲染兼容性。为解决这一差距,我们开展了首个针对AI生成网页跨环境兼容性的系统实证研究。具体而言,我们构建了WebCompat数据集,该数据集包含2032个带注释的实例,涵盖8种代表性AI工具生成的网页,每种网页均在9种浏览器-设备组合中进行渲染。我们分析了兼容性问题的普遍程度、用户可感知的症状以及底层代码级根本原因。研究结果显示,68%的生成网页存在至少一个兼容性问题,凸显了围绕MLLM生成的前端制品的普遍可靠性问题。最普遍的症状是破坏整个页面布局的故障(占88.3%):页面会直接缩小以适配目标屏幕,导致字体过小,或出现比例不匹配造成内容被截断;而局限于单个元素的故障,如图像失真或组件缺失,相对较少见(占13.4%)。此外,尽管大多数MLLM在生成过程中融入了响应式设计模式,但它们未能正确实现这些代码。基于上述发现,我们开发了XCompat,这是一种结合视觉截图与结构化DOM树进行分析的轻量级离线兼容性问题检测器,其在WebCompat测试集上的F1分数达到0.903,优于现有兼容性检查工具和LLM基线方法。所有数据集和工具均已发布,以支持未来基于MLLM的前端代码生成的渲染可靠性研究。

英文摘要

Multimodal Large Language Models (MLLMs) have been increasingly adopted to automate webpage generation from visual designs (e.g., screenshots). However, existing evaluations are limited to visual fidelity assessment under a fixed browser-device configuration. Such a setting overlooks the cross-environment rendering compatibility for real-world deployments. To address this gap, we present the first systematic empirical study of cross-environment compatibility in AI-generated webpages. Specifically, we construct WebCompat, a dataset of 2,032 annotated instances, comprising webpages generated by 8 representative AI tools, each rendered across 9 browser-and-device combinations. We analyze the prevalence of compatibility issues, their user-perceptible symptoms, and underlying code-level root causes. Our findings reveal that 68% of generated webpages exhibit at least one compatibility issue, underscoring the pervasive reliability concerns surrounding MLLM-generated front-end artifacts. The most prevalent symptoms are failures that disrupt the entire page layout (88.3%): pages shrink directly to fit the target screen with too small fonts, or exhibit scale mismatches that produce cut-off content. Failures localized to individual elements, such as image distortion or missing components, are comparatively less common (13.4%). Furthermore, although most MLLMs incorporate responsive design patterns into the generation, they fail to properly implement these codes. Guided by the findings, we develop XCompat, a lightweight offline compatibility issue detector that combines visual screenshots and the structural DOM tree for analysis. It achieves an F1 score of 0.903 on the WebCompat-test, outperforming the existing compatibility checking tools and LLM baselines. All datasets and tools are released to support future research on rendering reliability in MLLM-based front-end code generation.

↑