发表机构
University of California, Santa Cruz; Carnegie Mellon University(加州大学圣克鲁兹分校; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究小型语言模型和编码代理生成的网页前端代码违反界面质量原则的问题,通过统一多来源的界面质量原则,构建数据集并利用强化学习训练轻量级视觉语言模型作为评判器,提升检测效果并发布相关方法支持评估和后续工作。
AI 中文摘要
小型语言模型和编码代理越来越多地生成网页前端代码,但其输出通常主要根据功能正确性进行评估。生成的界面可能在编译、渲染并通过单元测试的情况下,仍违反既定的界面质量原则。现有审计方法在成本、覆盖范围和可扩展性之间存在权衡。我们研究轻量级视觉语言模型能否作为生成界面的有效评判器。我们统一了来自三个HCI知识互补来源的19条界面质量原则,并通过在干净的、由语言模型生成的Tailwind页面中合成注入已知违规来构建经过验证的数据集。通过在4B视觉语言模型上持续进行强化学习,微F1从36%提高到84%,19条原则中有13条超过80%的F1。所得评判器可审计生成的界面、过滤低质量界面训练数据并为设计感知代码生成提供奖励信号。我们还发布了数据生成方法和注入/验证提示,以支持可重复评估和未来关于可扩展界面质量评估的工作。
英文摘要
Small language models and coding agents increasingly generate web front-end code, yet their outputs are typically evaluated primarily for functional correctness. A generated interface may compile, render, and pass unit tests while still violating established interface quality principles, including accessibility barriers, deceptive design patterns, poor visual hierarchy, and excessive decision complexity. Existing auditing approaches face a trade-off between cost, coverage, and scalability: expert human review provides rich judgment but is slow and expensive; frontier vision-language models offer broader reasoning capabilities but remain costly to deploy at scale; and rule-based tools such as axe-core and Lighthouse are inexpensive but primarily capture mechanically checkable accessibility issues. We investigate whether a lightweight vision-language model can serve as an effective critic for generated interfaces. We unify 19 interface-quality principles from three complementary sources of HCI knowledge: WCAG 2.2 accessibility standards, deceptive design taxonomies, and established theories of perception, cognition, and interaction. To train this critic, we construct a verified dataset of approximately 10,000 generated web pages by synthetically injecting known violations into clean, LLM-generated Tailwind pages. Continued reinforcement learning on a 4B vision-language model improves micro-F1 from 36\% to 84\%, with 13 of 19 principles exceeding 80\% F1. The resulting critic can audit generated interfaces, filter low-quality interface training data, and provide a reward signal for design-aware code generation. We release our data-generation recipe and injection/verification prompts to support reproducible evaluation and future work on scalable interface-quality assessment.