EvoGenUI-Bench:评估大型语言模型作为多轮生成式UI助手的性能
EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants
浏览论文内容
中文总结 AI 辅助
本文提出EvoGenUI-Bench多轮生成式UI评估基准,发现现有8个大型语言模型在该基准上轮次通过率最高74.9%、五轮回合完成率仅37.3%,且工具接地任务的跨轮次保留率更低,揭示生成式UI需关注多轮状态同步问题。
中文摘要 AI 辅助
大型语言模型可生成交互式网页界面,但可靠的生成式UI需在用户请求演变过程中维持可执行的制品。本文提出EvoGenUI-Bench,这是一个用于多轮界面维护的基准,包含150个五轮任务,共750轮,涵盖信息展示、可执行交互、工具接地外部状态三种场景。我们在浏览器中执行生成的制品,并通过截图、源代码与DOM证据、行为者轨迹、运行时日志对其进行评估。除了轮次级和回合级成功指标外,我们还使用相邻通过保留率(Adjacent Pass Retention)衡量跨轮次保留能力。在8个模型中,即使是最强模型也仅达到74.9%的轮次通过率,且仅完成37.3%的五轮回合;在工具接地任务上,相邻通过保留率进一步降至52.4%。诊断分析显示,展示失败集中在信息架构,交互失败集中在派生状态传播和功能绑定,工具接地失败还涉及外部状态接地和需求分解。这些结果将生成式UI评估从孤立输出的判断重新定义为测试界面行为、派生状态、外部状态及助手声明是否随制品演变保持同步。
英文摘要
Large language models can generate interactive web interfaces, but reliable generative UI requires maintaining an executable artifact as user requests evolve. We introduce EvoGenUI-Bench, a benchmark for multi-turn interface maintenance comprising 150 five-turn tasks and 750 turns across three scenarios: information presentation, executable interaction, and tool-grounded external state. We execute generated artifacts in a browser and evaluate them using screenshots, source and DOM evidence, actor traces, and runtime logs. Beyond turn-level and episode-level success, we measure cross-turn retention with Adjacent Pass Retention. Across eight models, even the strongest achieves 74.9% Turn Pass while completing only 37.3% of five-turn episodes; APR further falls to 52.4% on tool-grounded tasks. Diagnostic analysis shows that presentation failures center on information architecture, interaction failures on derived-state propagation and affordance binding, and tool-grounded failures additionally involve external-state grounding and requirement decomposition. These results reframe generative UI evaluation from judging isolated outputs to testing whether interface behavior, derived state, external state, and assistant claims remain synchronized as the artifact evolves.
发表机构
- New York University Shanghai(上海纽约大学)
机构由 AI 辅助整理,请以论文原文为准。