arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越单一评判者:用于生成式UI评估的社会角色面板模拟

Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation

Zheng Wu, Yibo Luo, Pu Zhang, Cheng Yang, Zhuosheng Zhang

arXiv 2607.28439首次发表:更新:

AI 中文总结

该研究针对GenUI评估问题,提出三阶段ESPP方法,通过多样化角色面板提升评估保真度,揭示用户子群体的结构性分歧,代码已公开。

AI 中文摘要

生成式UI(GenUI)可让大型语言模型直接根据自然语言指令合成完整的、可渲染的界面,但其生成质量的评估仍是一个未解决的问题。人工评估成本高且评估者存在差异,而“LLM作为评判者”虽可扩展,但仅反映单一隐含观点,无法捕捉不同真实用户群体对同一界面的实际感知。我们提出基于证据的社会加权角色面板(ESPP),这是一种三阶段GenUI评估方法:由一组心理多样化、基于证据的角色独立对截图进行评分,在基于特质的语义门控有界置信机制下交换意见,再通过受Delphi启发的社会加权聚合为单一判断。ESPP与人工判断的匹配度远高于简单的单次评判者,将皮尔逊相关系数r从0.716提升至0.922,而提示集成对照组仅能弥补约三分之一的差距,表明角色与证据基础是提升的主要来源。除了保真度提升,保留每位小组成员的个人评分还发现,用户子群体在整体模型排名上达成一致,但在特定评分维度上存在显著分歧,这种结构性分歧是单一同质化评判者会系统性忽略的。代码可在此httpsURL获取。

英文摘要

Generative UI (GenUI) lets large language models synthesize a complete, renderable interface directly from a natural-language instruction, but evaluating the quality of what they generate remains an open problem. Human evaluation is costly and rater-variant, while LLM-as-a-judge is scalable but reflects only a single implicit viewpoint, unable to capture how different populations of real users actually perceive the same interface. We propose the Evidence-Grounded, Social-Weighted Persona Panel (ESPP), a three-stage GenUI evaluation method in which a panel of psychologically diverse, evidence-grounded personas independently rates a screenshot, exchanges opinions under a trait-derived, semantically-gated bounded-confidence mechanism, and is aggregated via Delphi-inspired social weighting into a single judgment. ESPP tracks human judgment substantially more closely than a naive single-pass judge, raising Pearson $r$ from $0.716$ to $0.922$, and a prompt-ensemble control recovers only about a third of this gap, isolating genuine persona and evidence grounding as the dominant source of improvement. Beyond this fidelity gain, retaining each panelist's individual rating further reveals that user subgroups agree on overall model rankings yet diverge sharply on specific rating dimensions, a structural disagreement a single homogeneous judge would systematically erase. The codes are available at https://github.com/Wuzheng02/ESPP.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑