arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05423cs.HCcs.CLcs.SE

看见而不理解:大型语言模型对移动用户界面质量、失败分类与架构解释的评估

Seeing Without Understanding: Large Language Model Evaluation of Mobile User Interface Quality, Failure Taxonomy, and Architectural Explanation

Md Rejaul Korim Sadi, Golam Mostofa Naeem, Toufiqur Rahman Tasin, Syed Mostofa Moosa, Mahmudul Hasan Emon, Mahmudur Rashid, Ferdus Ahmed

首次发表
浏览论文内容

中文总结 AI 辅助

本文利用RICO数据集构建评估语料库,通过启发式基线与多语言模型对比,系统检验LLM在移动UI质量评估中的可靠性,并归纳失败分类及架构解释。

中文摘要 AI 辅助

在软件工程和人机交互领域,大规模评估移动用户界面质量仍然是一个持续的挑战。基于规则的启发式方法提供了结构上的可靠性,但需要大量的工程投入,而人工标注则无法扩展到每年产生的应用程序数量。大型语言模型提供了一种有前景的替代方案,然而它们用于结构化界面判断的可靠性尚未得到系统性的检验,其失败背后的模式也仍未得到充分表征。本文旨在填补这两个空白。我们首先使用完整的RICO数据集,其中包含66,261个真实世界的移动应用屏幕,通过严格的、基于文献指导的筛选过程,从中构建了一个包含15,000个屏幕的精炼评估语料库。每个屏幕根据七个标准进行评估:结构化JSON有效性、最小可见元素数量、可点击组件存在性、非零布局边界、图像完整性以及感知重复移除。针对该语料库,我们应用了一个基于严重性加权可用性信号、归一化布局指标和像素比复杂度度量(这些度量已根据真实用户情感进行校准)的启发式基线。多个语言模型独立地对每个屏幕的可用性、布局质量和视觉复杂度进行评分,评分依据是结构化JSON描述和原始截图。与启发式方法的维度级比较使用了一致率、Cohen's Kappa和置信度校准。反复出现的分歧模式被组织成一个失败分类法,并通过Transformer架构特征进行解释:最大似然估计合理性偏差、注意力错位和自回归过度承诺。

英文摘要

Evaluating mobile user interface quality at scale remains a persistent challenge in software engineering and human-computer interaction. Rule-based heuristic methods offer structural reliability but demand significant engineering effort, while human annotation does not scale to the volume of applications produced annually. Large language models present a promising alternative, yet their reliability for structured UI judgment has not been systematically examined, and the patterns behind their failures remain insufficiently characterized. This paper addresses both gaps. We begin with the complete RICO dataset of 66,261 real-world mobile application screens, from which we derive a refined evaluation corpus of 15,000 screens through a rigorous, literature-guided selection process. Each screen is assessed across seven criteria: structural JSON validity, minimum visible element count, clickable component presence, non-zero layout bounds, image integrity, and perceptual duplicate removal. Against this corpus, we apply a heuristic baseline built from severity-weighted usability signals, normalized layout metrics, and pixel-ratio complexity measures calibrated to real user sentiment. Multiple language models independently rate each screen across usability, layout quality, and visual complexity from structured JSON descriptions and raw screenshots. Dimension-level comparison against the heuristic uses agreement rates, Cohen's Kappa, and confidence calibration. Recurring divergence patterns are organized into a failure taxonomy and interpreted through transformer architectural signatures: MLE plausibility bias, attention misgrounding, and autoregressive over-commitment.

发表机构

  • Metropolitan University(大都会大学)

机构由 AI 辅助整理,请以论文原文为准。

↑