arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09654cs.AI

无幻觉的GUI定位:基于无回归的布局感知匹配

Hallucination-Free GUI Grounding via Regression-Free Layout-Aware Matching

Yuke Li, Xuehan Hou

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出无回归框架,通过解耦指令理解与布局感知定位,在ScreenSpot-Pro和Mind2Web数据集上显著提升GUI定位的准确率、成功率及元素选择率,抑制了坐标幻觉。

中文摘要 AI 辅助

GUI智能体正从依赖元数据的大语言模型转向直接基于截图运行的纯视觉多模态大语言模型(MLLM)。核心任务GUI定位要求将抽象用户指令转化为精确的元素坐标,该任务面临双重长期障碍:传统定位模型缺乏解读抽象指令的语义丰富性,而端到端MLLM则因细粒度感知不足出现坐标幻觉。我们提出一种无回归框架:冻结的MLLM负责指令解析,专用定位模型处理精确定位,无需学习任何坐标回归。冻结的MLLM首先将抽象指令细化为包含丰富布局线索的结构化视觉描述,这些描述随后输入新型布局感知GUI定位模型,该模型通过与布局先验候选匹配执行无回归定位,从本质上抑制幻觉并避免昂贵的微调。定位模型仅用文本/图标二元标签训练,无需坐标回归参数。在ScreenSpot-Pro上,我们的方法相比端到端系统实现了20%以上的定位准确率提升;在Mind2Web上,它将成功率和元素选择率提高了15%以上。这些结果表明,将指令理解与布局感知定位解耦可有效解决GUI交互的核心挑战。

英文摘要

GUI agents are shifting from metadata-dependent large language models to purely visual multimodal large language models (MLLMs) that operate directly on screenshots. The core task, GUI grounding, requires translating abstract user instructions into precise element coordinates. This task faces a persistent dual obstacle: conventional grounding models lack the semantic richness to interpret abstract instructions, while end-to-end MLLMs suffer from coordinate hallucinations caused by deficient fine-grained perception. We propose a regression-free framework where a frozen MLLM performs instruction parsing and a dedicated grounding model handles precise localization without learning any coordinate regression. A frozen MLLM first elaborates the abstract instruction into a structured visual description rich in layout cues. These descriptions are then fed to a novel Layout-Aware GUI Grounding Model, which performs regression-free localization by matching against layout-prior candidates, inherently suppressing hallucinations and avoiding expensive fine-tuning. The grounding model is trained with only Text/Icon binary labels, requiring no coordinate regression parameters. On ScreenSpot-Pro, our method achieves over 20% improvement in grounding accuracy over end-to-end systems; on Mind2Web, it raises success rate and element selection rate by more than 15%. These results demonstrate that decoupling instruction understanding from layout-aware localization effectively resolves the core challenges of GUI interaction.

发表机构

  • School of Electronic and Computer Engineering, Peking University(北京大学电子与计算机工程学院)

机构由 AI 辅助整理,请以论文原文为准。

↑