arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

循环潜在视觉搜索用于GUI接地

Recurrent Latent Visual Search for GUI Grounding

Kaiyu Wu, Beichen Zheng, Weiyao Huang, Keze Wang

arXiv 2610.05185首次发表:更新:

发表机构

Sun Yat-sen University(中山大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出ReLaViS,一种在单轮交互中通过循环潜在视觉搜索实现显式空间多步搜索的GUI接地方法,基于Qwen2.5-VL-7B,在ScreenSpot-Pro上准确率提升3.1个百分点至56.3%,推理开销仅增3.5%,并在五个基准上均优于单步基线。

AI 中文摘要

GUI接地是视觉语言模型驱动的GUI智能体的关键能力,帮助它们通过定位截图中的相应元素来执行用户指令。单步接地在处理小元素和密集布局时存在困难,这推动了多步视觉搜索的发展。然而,现有方法通常依赖于与视觉空间不对齐的文本推理,或与外部视觉工具进行昂贵的多轮交互。为了使多步视觉搜索成为模型内的显式空间过程,我们提出了ReLaViS,它在单次交互轮次中执行循环潜在视觉搜索。在每一步中,空间搜索头使用隐藏状态查询截图的视觉令牌,生成显式表示搜索焦点的空间搜索分布。该分布随后将视觉令牌聚合为潜在视觉证据,并循环反馈作为下一个输入嵌入以调节后续搜索。我们进一步通过从扁平元素注释构建的轨迹引入GUI感知的从粗到细的归纳偏置,从全局界面通过中间元素组监督搜索直至目标。基于Qwen2.5-VL-7B,ReLaViS将ScreenSpot-Pro准确率提高了3.1个百分点达到56.3%,推理FLOPs仅增加3.5%,并在所有五个基准上优于匹配的单步基线。

英文摘要

GUI grounding is a critical capability for GUI agents powered by vision-language models, helping them execute user instructions by locating the corresponding elements in screenshots. Single-step grounding struggles with small elements and dense layouts, motivating multi-step visual search. However, existing approaches commonly rely on textual reasoning misaligned with visual space or costly multi-round interactions with external visual tools. To make multi-step visual search an explicit spatial process within the model, we propose ReLaViS, which performs Recurrent Latent Visual Search in a single interaction round. At each step, a spatial search head uses the hidden state to query the screenshot's visual tokens, producing a spatial search distribution that explicitly represents the search focus. This distribution then aggregates the visual tokens into latent visual evidence, which is recurrently fed back as the next input embedding to condition subsequent search. We further introduce a GUI-aware coarse-to-fine inductive bias through trajectories constructed from flat element annotations, supervising search from the global interface through intermediate element groups to the target. Built on Qwen2.5-VL-7B, ReLaViS improves ScreenSpot-Pro accuracy by 3.1 percentage points to 56.3% with only a 3.5% increase in inference FLOPs and outperforms the matched single-step baseline on all five benchmarks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑