GUI-Lens:面向通用视觉语言模型(VLM)的GUI定位粗到精裁剪框架
GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs
浏览论文内容
中文总结 AI 辅助
GUI-Lens是面向通用VLM的GUI定位粗到精裁剪框架,通过主动视觉观察逐步聚焦确定目标,在四个基准和三个VLM后端上使定位准确率最高提升24.9个百分点,在GPT-5.5上达SOTA性能。
中文摘要 AI 辅助
GUI定位任务是将自然语言指令映射到点击位置,是可靠GUI智能体的核心,但在高分辨率、控件密集的界面上仍具挑战性,因为视觉语言模型(VLM)可能识别出目标控件却无法精确定位到可交互程度。现有多数方法提供各类定位辅助,但仍依赖直接点击预测,会让视觉歧义或不准确的初始估计传递到最终结果。本文提出GUI-Lens,一种允许通用VLM通过主动视觉观察确定目标的粗到精定位框架。具体而言,GUI-Lens从截图中提取OCR文本和检测到的UI组件,并将其位置作为坐标参考;VLM利用指令、当前视图及这些参考,选择下一个视图的区域和尺度,对其进行裁剪并放大以提供更精细的视觉细节;该过程会在连续聚焦的视图上持续进行,直到确定目标。整个过程中会检查提出的裁剪和点击是否符合指令,最终将局部位置映射回原始屏幕坐标。在四个GUI定位基准和三个通用VLM后端上的实验表明,GUI-Lens使整体定位准确率最高提升24.9个百分点,且在GPT-5.5后端上达到了最先进的性能。
英文摘要
GUI grounding maps natural-language instructions to click locations and is essential for reliable GUI agents. The task remains difficult on high-resolution, densely populated interfaces because a vision-language model (VLM) may recognize a requested control without locating it precisely enough for interaction. Most existing methods provide various forms of localization assistance, but still rely on a direct click prediction, allowing visual ambiguity or an inaccurate initial estimate to propagate to the final result. In this paper, we introduce GUI-Lens, a coarse-to-fine grounding framework that allows a general-purpose VLM to determine the target through active visual observations. Specifically, GUI-Lens extracts OCR text and detected UI components from the screenshot and presents their positions as coordinate references. Using the instruction, the current view, and these references, the VLM selects the region and scale of the next view, which is cropped and enlarged to provide finer visual details. This process continues over successively focused views until the target is determined. Proposed crops and clicks are checked against the instruction throughout the process, and the final local position is mapped back to the original screen coordinates. Experiments on four GUI grounding benchmarks and three general-purpose VLM backends show that GUI-Lens improves overall grounding accuracy by up to 24.9 percentage points and achieves state-of-the-art performance with GPT-5.5.