发表机构
University of Southern California; USC Information Sciences Institute; Qatar Computing Research Institute (QCRI)(南加州大学; 南加州大学信息科学研究所; 卡塔尔计算研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出GUI-Primitives基准,通过994个含7种空间关系的对比指令对,发现多数视觉-语言GUI定位模型的失败源于候选定位而非关系理解,且标记候选可提升准确率。
AI 中文摘要
计算机使用智能体将自然语言指令与截图关联以定位界面元素,但现有基准无法区分模型是否将关系语言绑定到了正确元素。本文提出GUI-Primitives,这是一个包含994个对比指令对的基准,涵盖图形用户界面中的7种空间关系(左/右、上/下、包含、对齐、邻近、列表序数、遮挡)。每对指令保持截图和锚点固定,仅改变关系表达式,使正确目标在两个指定候选间切换。5名标注员验证了196个样本子集(一致性κ=0.94,结构合理性;κ=0.79,目标选择)。19个视觉-语言模型的严格框内准确率最高达32%。由于模型输出无约束坐标,本文按预测坐标所属的候选区域对每个预测分类:60%-92%的样本预测落在两个候选之外。在落在候选区域内的条件下,水平位置、垂直位置、邻近和列表序数的目标选择准确率达0.82-0.90,但包含和遮挡的准确率与0.50无显著差异:多数失败源于候选定位而非关系理解。10个模型的基准准确率与ScreenSpot-Pro准确率呈正相关(斯皮尔曼ρ=+0.74),这是该样本量下的探索性关联。标记两个指定候选可使选择准确率提升35-57个百分点,这是一种提供候选集而非可部署方法的 oracle 诊断。本文发布该基准、预测结果和代码。
英文摘要
Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether models bind relational language to the correct element. We introduce GUI-Primitives, a 994-item benchmark of contrastive instruction pairs over seven spatial relations in graphical user interfaces (left/right, above/below, containment, alignment, proximity, list ordinal, occlusion). Each pair holds the screenshot and anchor fixed while changing the relation expression, so the correct target moves between two designated candidates. Five annotators validate a 196-item subset ($κ= 0.94$ well-formedness; $κ= 0.79$ target selection). Nineteen vision-language models reach at most $32\%$ strict point-in-box accuracy. Because models emit unconstrained coordinates, we classify each prediction by the candidate region it falls within. Predictions fall outside both candidates on $60-92\%$ of items. Conditional on falling within a candidate region, target selection reaches 0.82-0.90 for horizontal position, vertical position, proximity, and list ordinal, but does not differ significantly from 0.50 for containment and occlusion: most failures reflect candidate localization rather than relation understanding. Across ten models, benchmark accuracy correlates with ScreenSpot-Pro accuracy (Spearman $ρ= +0.74$), an exploratory association at this sample size. Marking the two designated candidates raises selection accuracy by 35--57 percentage points, an oracle diagnostic that supplies the candidate set rather than a deployable method. We release the benchmark, predictions, and code.
CommentsAccepted to EMNLP 2026 Main Conference