发表机构
University of Helsinki(赫尔辛基大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对GUI定位评估中嵌入相似度与语义定位的混淆问题,对比编码器与词汇基线,提出需报告词汇基线等诊断指标以区分标签恢复与语义定位。
AI 中文摘要
将UI元素作为文本元数据处理的GUI定位评估,常将高指令-元素嵌入相似度视为语义定位的证据。在三个移动端与网页基准测试中,我们表明该解释常被可见标签恢复混淆:词汇基线在top-1表现仍具竞争力,仅文本方法在标签匮乏目标上表现较弱,编码器top-1命中可通过词汇排名、候选池大小及标签类型预测。我们将每个动作评估为同屏幕排序任务,对比5种现成单向量编码器与词汇基线。编码器可恢复部分词汇遗漏,但可部署的融合增益远小于目标感知的神谕增益。这些发现表明,基于嵌入的评估可能将可见标签恢复与语义GUI定位混淆,因此应报告词汇基线、标签类型分层及可部署融合诊断。我们发布的代码库提供分析脚本及去文本化的逐步面板:this https URL。
英文摘要
GUI grounding evaluations that expose UI elements as text metadata often treat high instruction-element embedding similarity as evidence of semantic grounding. Across three mobile and web benchmarks, we show that this interpretation is frequently confounded by visible-label recovery. Lexical baselines remain competitive at top-1, label-poor targets remain weak for text-only methods, and encoder top-1 hits are predictable from lexical rank, candidate-pool size, and label type. We evaluate each action as a same-screen ranking task, comparing five off-the-shelf single-vector encoders with lexical baselines. Encoders recover some lexical misses, but deployable fusion gains are much smaller than target-aware oracle gains. These findings show that embedding-based evaluations can conflate visible-label recovery with semantic GUI grounding. Embedding-based evaluations should therefore report lexical baselines, label-type stratification, and deployable-fusion diagnostics. Our released repository provides analysis scripts and detexted per-step panels: https://github.com/qijia123/lexical-coupling-release.