发表机构
The Hong Kong University of Science and Technology (Guangzhou); Alibaba Group; The Hong Kong University of Science and Technology(香港科技大学(广州); 阿里巴巴集团; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
AUV-Bench通过1,395个Web界面和四项任务评估多模态模型的美学能力,揭示模型在评分上中等一致,但诊断成功率低且存在判断-行动差距,表明其缺乏细粒度美学理解。
AI 中文摘要
多模态基础模型越来越多地被用于评估和生成用户界面(UI),通常能产生看似合理的美学判断和视觉上可行的页面。然而,在专业设计审查下,它们的行为可能与人类设计师的行为存在显著差异。在专业设计实践中,设计师依赖一套系统的美学原则,这些原则始终如一地指导判断、诊断、修复和创作。因此,连贯的美学能力应将美学判断与设计行动联系起来。然而,现有的评估通常孤立地评估这些能力,使得难以确定任务层面的成功是否反映了共享的美学理解,还是仅仅是零散的任务特定能力。为解决这一差距,我们引入了AUV-Bench,与专业UI设计师合作开发,围绕1,395个可执行的Web界面和四项任务:美学评分、诊断、修复和文本到UI生成。这些任务共享一组UI和美学原则,诊断和修复进一步在660个受控退化实例上对齐,以实现对判断和行动的实例级分析。对12个模型的评估揭示了一种能力不平衡:模型在整体美学评分上与专业设计师表现出中等程度的一致性,但精确诊断链的成功率最高仅为24.7%。在对齐的诊断-修复案例中,正确的判断和成功的修复并不总是同时发生,暴露了识别美学问题与成功采取行动之间的判断-行动差距。在开放式生成中,即使领先的模型在人类校准的评估下也只能达到中等的美学质量。总体而言,当前模型表现出部分美学能力,但仍缺乏可靠UI设计所需的细粒度理解和判断-行动连贯性。
英文摘要
Multimodal foundation models are increasingly used for evaluating and generating user interfaces (UIs), often producing seemingly reasonable aesthetic judgments and visually plausible pages. However, under professional design scrutiny, their behavior can differ substantially from that of human designers. In professional design practice, designers rely on a systematic set of aesthetic principles that consistently guide judgment, diagnosis, repair, and creation. A coherent aesthetic capability should therefore connect aesthetic judgment with design actions. Existing evaluations, however, typically assess these abilities in isolation, making it difficult to determine whether task-level success reflects a shared aesthetic understanding or merely fragmented task-specific competence. To address this gap, we introduce AUV-Bench, developed in collaboration with professional UI designers around 1,395 executable web interfaces and four tasks: aesthetic scoring, diagnosis, repair, and text-to-UI generation. The tasks share a pool of UIs and aesthetic principles, with diagnosis and repair further aligned on 660 controlled-degradation instances to enable instance-level analysis of judgment and action. Evaluation of 12 models reveals a capability imbalance: models show moderate agreement with professional designers in holistic aesthetic scoring, yet exact diagnosis-chain success peaks at only 24.7%. On the aligned diagnosis-repair cases, correct judgments and successful repairs do not consistently coincide, exposing a Judgment-Action Gap between identifying aesthetic problems and successfully acting on them. In open-ended generation, even leading models achieve only moderate aesthetic quality under human-calibrated evaluation. Overall, current models exhibit partial aesthetic competence, but still lack the fine-grained understanding and judgment-action coherence required for reliable UI design.