发表机构
Kombai Inc.(Kombai公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出Design Creativity Bench基准,发现LLM生成UI设计的原创性、创意范围显著低于人类,恰当性略高于人类,但重复性远高于人类,需针对性改进。
AI 中文摘要
随着领先的大语言模型(LLM)在能力评估上不断进步,它们在设计任务中生成创意输出时的局限性仍未得到充分表征。本研究推出Design Creativity Bench,这一基准用于评估用户界面(UI)设计的多样性与恰当性。它测量同一提示下不同模型设计的独特性(原创性)、同一UI目标在不同产品领域的两个提示间模型设计的变化程度(创意范围),以及每个设计满足需求文档中接受标准的比例(恰当性)。不同模型的同提示设计对的原创性为0.592(95%置信区间[0.582, 0.602]),远低于同提示人机设计对的0.764(95%置信区间[0.751, 0.778]);模型间的创意范围为0.581(95%置信区间[0.567, 0.597]),而人类设计的创意范围为0.902(95%置信区间[0.884, 0.919]);所有模型的恰当性均超过90%,最优模型达到99.2%,略高于人类设计的98.0%。本研究表明,LLM的默认输出虽总体恰当,但相比人类基准存在明显的重复性问题,亟需采取有力措施解决该问题。
英文摘要
As leading LLMs improve on capability evaluations, their limitations in producing creative outputs on design tasks remain insufficiently characterised. Our work introduces Design Creativity Bench, a benchmark that evaluates diversity and appropriateness in UI designs. It measures distinctiveness among models on the same prompt (originality), how much a model's designs change between two prompts for the same UI goal in different product domains (creative range), and the share of a brief's acceptance criteria each design meets (appropriateness). Originality is 0.592 for same-prompt design pairs from different models (95% CI [0.582, 0.602]), far below the 0.764 for same-prompt human-model pairs (95% CI [0.751, 0.778]). Creative range is 0.581 across models (95% CI [0.567, 0.597]), against 0.902 for human designs (95% CI [0.884, 0.919]). Appropriateness is above 90% for every model, and the best model reaches 99.2%, slightly above the 98.0% for human designs. Our work shows that the default output of LLMs, though generally appropriate, is substantially more repetitive than the human baseline. This calls for strong measures to address the issue.