AI 中文总结
本文针对现有GUI定位方法在小目标等场景性能骤降的问题,提出LookAgain闭环GUI定位器,通过预测后视觉反思的多轮过程结合SFT与GRPO训练,在相关基准上达到SOTA性能。
AI 中文摘要
近期的图形用户界面(GUI)定位器在标准基准测试上的单样本准确率已显著提升,但在小目标、密集排列控件及分布外界面上的性能会急剧下降。我们将该差距归因于现有方法共有的范式性局限:它们均未将生成的坐标视为需在新视觉证据下反思和修正的假设,这体现为三个耦合问题:1)缺乏事后反思:预测结果在生成时即被冻结,无内部机制对其进行质疑或修正;2)视觉证据与预测解耦:辅助视觉证据被用于支持后续坐标,而非审视已确定的坐标;3)基于视图而非预测进行修正:迭代放大仅修正被检查的区域,而非继承前序坐标作为需修正的空间先验。本文提出LookAgain,一种由预测后视觉反思驱动的闭环GUI定位器。LookAgain将定位重新表述为多轮「预测-再看-修正」过程,包含两个核心原语:「定位」生成坐标假设,在图像上渲染标记并附加预测区域的局部补丁,将后续推理步骤锚定在前序预测作为空间先验;「确认」接受或拒绝该假设并终止流程。我们在构建的反思轨迹上对LookAgain定位器进行SFT作为冷启动,随后以终端定位正确性为唯一奖励进行GRPO训练。大量实验表明,LookAgain在拒绝感知型和通用GUI定位基准上均实现了性能提升,达到了最先进的结果,全面的 ablation 进一步验证了所提框架的有效性。
英文摘要
Recent graphical user interface (GUI) grounders have significantly advanced single-shot accuracy on standard benchmarks, yet their performance degrades sharply on small targets, densely packed controls and out-of-distribution interfaces. We attribute this gap to a paradigmatic limitation shared by existing approaches: none of them treats a produced coordinate as a hypothesis to be reflected upon and revised under new visual evidence. This manifests as three coupled issues: 1) Lack of post-hoc reflection. The prediction is frozen at the moment of emission, leaving no internal mechanism to challenge or refine it. 2) Visual evidence decoupled from the prediction. The auxiliary visual evidence is gathered to support the upcoming coordinate rather than to scrutinise the one already committed to. 3) Refinement over views, not over predictions. The iterative zoom-in refines the inspected region instead of inheriting a previous coordinate as a spatial prior to be corrected. In this paper, we propose LookAgain, a closed-loop GUI grounder driven by post-prediction visual reflection. LookAgain reformulates grounding as a multi-turn predict-look-again-refine process with two primitives: "locate" posts a coordinate hypothesis, renders a marker on the image and appends a local patch of the predicted region. It anchors the next reasoning step to the previous prediction as a spatial prior; "confirm" accepts or reject the hypothesis and terminates the procedure. We train the LookAgain grounder with SFT on constructed reflective trajectories as a cold start, followed by GRPO with terminal grounding correctness as the sole reward. Extensive experiments show that LookAgain consistently improves performance on both refusal-aware and general GUI grounding benchmarks, achieving state-of-the-art results. Comprehensive ablations further verify the effectiveness of the proposed framework.