发表机构
Shanghai University of Engineering Science; Fudan University(上海工程技术大学; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM的GUI生成评估难题,提出SchemaGUI基准,测试Qwen3.5等5种主流模型,发现几何控制、布局复杂度、思维模式对GUI生成效果的影响规律。
AI 中文摘要
大型语言模型(LLM)在图形用户界面(GUI)生成方面展现出强大潜力,但由于数据分布不受控、标注存在噪声以及布局场景覆盖有限,可靠的评估工作仍颇具挑战。为解决该问题,我们提出SchemaGUI,这是一种基于模板的可控GUI生成评估基准。通过从参数化界面模式中合成配对的自然语言指令与确定性函数调用参考,SchemaGUI可在数秒内生成数千个带有确定性标注的任务,且无需人工标注。基于6种代表性双语场景中每种场景及语言各1000个评估实例,我们对5种主流模型进行了基准测试,包括Qwen3.5系列、Qwen3-Coder-30B以及DeepSeek-R1。我们的广泛分析揭示了3项关键见解:第一,精确的几何空间控制仍是重要瓶颈;将Qwen3.5从4B扩展至27B时,模式可行性从91.56%提升至99.63%,但几何分数提升幅度较小(从67.05%升至75.30%)。第二,生成难度对布局复杂度高度敏感,当前LLM擅长简单的序列排列,但在密集网格与多区域组合中会出现严重的坐标漂移。第三,思维模式会增加令牌消耗,同时通常会降低GUI分数,尤其对于较小规模的模型。
英文摘要
Large language models (LLMs) have demonstrated strong potential in graphical user interface (GUI) generation, but reliable evaluation remains challenging due to uncontrolled data distributions, noisy annotations, and limited layout scenario coverage. To address this, we propose SchemaGUI, a template-based benchmark for controllable GUI generation evaluation. By synthesizing paired natural language instructions and deterministic function-call references from parameterized interface schemas, SchemaGUI can generate thousands of deterministically annotated tasks in seconds without human labeling. Based on 1,000 evaluated instances per scenario and language across six representative bilingual scenarios, we benchmark five mainstream models, including the Qwen3.5 family, Qwen3-Coder-30B, and DeepSeek-R1. Our extensive analysis reveals three key insights. First, precise geometric spatial control remains an important bottleneck; while scaling Qwen3.5 from 4B to 27B improves Schema Feasibility from 91.56% to 99.63%, the Geometry score improves more modestly (from 67.05% to 75.30%). Second, generation difficulty is highly sensitive to layout complexity, with current LLMs excelling at simple sequential arrangements but suffering severe coordinate drift in dense grids and multi-region compositions. Third, thinking mode increases token consumption while generally reducing GUI Score, particularly for smaller models.
Comments18 pages, 6 figures, and 7 tables. Code is available at https://github.com/xdong2002/SchemaGUI