发表机构
Vercel; Claude Artifacts; ChatGPT Canvas; Firebase Studio; Bolt; Anthropic; OpenAI; Google; StackBlitz(Vercel公司; Claude Artifacts公司; ChatGPT Canvas公司; Firebase Studio公司; Bolt公司; Anthropic公司; OpenAI公司; 谷歌公司; StackBlitz公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究生成式UI工具中设计原理与实际界面的脱节现象(设计剧场),引入含24个任务的基准及指标,评估五个工具创建的120个界面,发现原理实现率低、工具识别原则有限,贡献概念、基准及评估结果。
AI 中文摘要
生成式用户界面工具有望通过将自然语言描述转化为完整界面来实现用户界面设计的民主化。除了界面,这些工具还会生成面向用户的设计原理来解释其布局、可访问性和设计选择。然而,尚不清楚这些陈述的原理是否实际反映在它们生成的界面中。我们将这种脱节称为“设计剧场”:看似合理且自信的设计原理与实际实现几乎没有关系。为了研究这一现象,我们引入了一个基准和三个衡量设计剧场的指标。该基准包括24个跨越结构、样式和功能设计要求的用户界面生成任务。使用这个基准,我们评估了由五个生成式用户界面工具创建的120个界面。平均而言,超过25%的面向用户的设计原理在生成的界面中未得到实现,对于功能要求,实现失败率增加到34%。工具识别出提示中嵌入的大约一半用户体验原则(平均值 = 0.54),五个工具中有四个实现的功能原则不到6%。我们还测量了不同工具之间的界面相似度,发现视觉外观和布局组织存在趋同,颜色选择的差异更大。总体而言,我们贡献了:1)设计剧场的概念;2)一个带有指标的基准,用于评估生成式用户界面工具陈述的推理是否反映在其实现中;3)对这些工具进行系统评估的结果。我们讨论了这些发现对生成式用户界面工具的设计和评估意味着什么。
英文摘要
Generative UI tools promise to democratize UI design by turning natural language descriptions into complete interfaces. Alongside the interface, these tools generate user-facing design rationales that explain their layout, accessibility, and design choices. However, it remains unclear whether these stated rationales are actually reflected in the interfaces they produce. We call this disconnect ``Design Theater'': plausible and confident design rationales that have little relationship to the actual implementation. To study this phenomenon, we introduce a benchmark and three metrics for measuring Design Theater. The benchmark includes 24 UI generation tasks spanning structural, styling, and functional design requirements. Using this benchmark, we evaluate 120 interfaces created by five generative UI tools. On average, over 25\% of user-facing design rationales are not implemented in the generated interface, and the implementation failure increases to 34\% for functional requirements. Tools recognize roughly half of the UX principles embedded in prompts (mean = 0.54), with four of five tools implementing 6\% or fewer functional principles. We also measure interface similarity across tools and find convergence in visual appearance and layout organization, with greater variation in color choices. Overall, we contribute: 1) the concept of Design Theater; 2) a benchmark with metrics for assessing whether the stated reasoning of generative UI tools is reflected in their implementations; 3) and findings from a systematic evaluation of these tools. We discuss what these findings mean for the design and evaluation of generative UI tools.
CommentsAccepted at AAAI/AIES