arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18690cs.CVcs.CLcs.HC

RankGround:基于轻量级重排序器引导裁剪选择的高效高分辨率GUI定位

RankGround: Efficient High-Resolution GUI Grounding via Lightweight Reranker-Guided Crop Selection

Liyang Fan, Xinping Bi, Yitai Li, Shuaimin Li, Hui Li, Min Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对GUI定位中精度与效率的权衡,提出RankGround两阶段框架,利用轻量级重排序器GroundRanker选择最优裁剪,实现单次VLM调用,推理加速1.4倍且平均精度提升5.5%。

中文摘要 AI 辅助

图形用户界面(GUI)定位是多模态智能体的基础感知任务,使其能够解释自然语言指令并与数字界面交互。现有方法在准确性和效率之间面临根本性权衡:直接的全图像推理通常无法捕获较小或视觉上相似的UI元素,而多裁剪策略虽然提高了定位精度,但代价是每次查询需要多次昂贵的视觉语言模型(VLM)调用。为解决这一挑战,我们提出了RankGround,一个两阶段框架,通过每次查询仅一次VLM调用即可实现准确的GUI定位。我们方法的核心是GroundRanker,一个轻量级多模态重排序器,能够从密集的候选集中识别最有希望的裁剪区域。由于没有现成的排序数据集,我们从现有的定位数据集构建排序监督数据。严格的包含标准和边界感知的正样本增强改善了杂乱布局中的对齐和空间覆盖。随后,GroundRanker通过两阶段课程进行训练:逐点目标首先学习粗略的包含关系,列表式目标则细化视觉相似裁剪之间的细微语义和空间差异。实验结果表明,RankGround在降低计算成本的同时,持续优于强基线方法。在所有骨干网络和屏幕尺度上,它实现了1.4倍的推理加速,并且平均定位精度比第二好的方法提高了5.5%,在GUI定位的效率和精度方面均确立了新的最先进水平。

英文摘要

Graphical User Interface (GUI) grounding is a fundamental perception task for multimodal agents, enabling them to interpret natural language instructions and interact with digital interfaces. Existing methods face a fundamental trade-off between accuracy and efficiency: direct full-image inference often fails to capture small or visually similar UI elements, while multi-crop strategies improve localization at the cost of multiple expensive Vision-Language Model (VLM) calls per query. To address this challenge, we propose RankGround, a two-stage framework that achieves accurate GUI grounding with a single VLM call per query. Central to our approach is GroundRanker, a lightweight multimodal reranker that identifies the most promising crop from a dense candidate set. Because no off-the-shelf ranking dataset is available, we construct ranking supervision data from existing grounding datasets. A strict containment criterion and boundary-aware positive augmentation improve alignment and spatial coverage in cluttered layouts. GroundRanker is then trained with a two-stage curriculum: a pointwise objective first learns coarse containment, and a listwise objective refines subtle semantic and spatial distinctions among visually similar crops. Experimental results show that RankGround consistently outperforms strong baselines while reducing computational cost. It achieves 1.4 times faster inference and improves localization accuracy by 5.5% on average over the second-best method across all backbones and screen scales, establishing a new state of the art in both efficiency and precision for GUI grounding.

发表机构

  • Shenzhen University of Advanced Technology(深圳理工大学)
  • Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences(中国科学院深圳先进技术研究院)
  • Xiamen University(厦门大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑