AI 中文总结
ScreenHaystack基准发现GUI接地模型存在空间盲区,并证明其源于训练数据覆盖不均,提出盲区导向增强策略以提升精度。
AI 中文摘要
我们引入了ScreenHaystack,一个用于评估GUI接地中空间可靠性的动态针-在-干草堆基准。不同于在固定位置测试每个目标,ScreenHaystack系统性地将受控目标图标重新定位到高分辨率GUI背景中,并测量模型能否一致地定位它们。使用此基准,我们发现领先的GUI接地模型,包括Qwen3-VL、UI-TARS、GTA和UI-Venus,表现出盲区:即尽管目标外观和指令固定,接地精度仍急剧下降的空间区域。这些盲区可转移到未见过的ScreenSpot-Pro示例中:盲区内的目标始终更难接地,Qwen3-VL-8B下降了16.1个百分点,而受控重定位显示,将目标移入盲区会降低精度,移出则提高精度。我们进一步通过受控合成实验表明,训练数据中不均匀的空间覆盖可诱发此类盲区。因此,我们提出一种简单策略,即盲区导向增强,它在盲区中添加监督,并在ScreenSpot-Pro上相比原始和随机增强的Click-100k微调提高了精度。
英文摘要
We introduce ScreenHaystack, a dynamic needle-in-a-haystack benchmark for evaluating spatial reliability in GUI grounding. Instead of testing each target at a fixed position, ScreenHaystack systematically relocates controlled target icons across high-resolution GUI backgrounds and measures whether models can localize them consistently. Using this benchmark, we find that leading GUI grounding models, including Qwen3-VL, UI-TARS, GTA, and UI-Venus, exhibit blind zones: spatial regions where grounding accuracy drops sharply despite fixed target appearance and instruction. These blind zones transfer to unseen ScreenSpot-Pro examples: targets inside blind zones are consistently harder to ground, with Qwen3-VL-8B dropping by 16.1 percentage points, and controlled relocation shows that moving targets into blind zones decreases accuracy while moving them out improves accuracy. We further show through controlled synthetic experiments that uneven spatial coverage in training data can induce such blind zones. Therefore, we propose a simple strategy, blind-zone-oriented augmentation, which adds supervision in blind zones and improves ScreenSpot-Pro accuracy over both original and randomly augmented Click-100k fine-tuning.
CommentsAccepted to EMNLP 2026 Main Conference