从可靠负样本中学习:基于置信度锚定的测试时自适应用于GUI接地
Learning from Reliable Negatives: Confidence-Anchored Test-Time Adaptation for GUI Grounding
浏览论文内容
中文总结 AI 辅助
本文提出基于坐标标记置信度的无标签测试时训练方法CANL,利用负样本学习提升GUI接地性能,在ScreenSpot-Pro上较基础模型提升8.9%。
中文摘要 AI 辅助
图形用户界面(GUI)接地对于自主智能体将自然语言指令映射到精确的屏幕坐标至关重要。然而,现有的监督微调和强化学习方法受到高标注成本的制约,造成了可扩展性瓶颈。在本文中,我们引入了一种无标签的测试时训练范式,其驱动力来自两个关键洞察:(1)坐标标记中的置信度模式比全序列置信度是更好的指标;(2)在稀疏的GUI坐标空间中,负样本比可能带有噪声的正样本提供更可靠的学习信号。我们首先提出置信度锚定学习(CAL),它利用坐标标记置信度来过滤伪标签并分配基于距离的二元奖励。在此基础上,我们开发了置信度锚定负样本学习(CANL),该方法仅使用负样本优化模型,以规避错误正样本的风险。实验结果表明,CANL-7B在ScreenSpot-V2上达到92.1%的准确率。在更具挑战性的ScreenSpot-Pro上,CANL-7B达到33.8%,相比基础模型有8.9%的绝对提升。我们的发现确立了坐标标记置信度作为手动标注的可扩展替代方案,用于可扩展的GUI智能体开发。
英文摘要
Graphical User Interface (GUI) grounding is essential for autonomous agents to map natural language instructions to precise screen coordinates. However, existing supervised fine-tuning and reinforcement learning methods are constrained by the high cost of annotation, creating a scalability bottleneck. In this paper, we introduce a label-free test-time training paradigm driven by two key insights: (1) confidence patterns in coordinate tokens are a better indicator than full-sequence confidence, and (2) in sparse GUI coordinate spaces, negative samples offer more reliable learning signals than potentially noisy positive ones. We first propose Confidence-Anchored Learning (CAL), which utilizes coordinate-token confidence to filter pseudo-labels and assign distance-based binary rewards. Building on this, we develop Confidence-Anchored Negative Learning (CANL), which exclusively optimizes the model using negative samples to bypass the risks of incorrect positive samples. Experimental results demonstrate that CANL-7B achieves 92.1% on ScreenSpot-V2. On more challenging ScreenSpot-Pro, CANL-7B reaches 33.8%, an 8.9% absolute improvement over the base model. Our findings establish coordinate-token confidence as a powerful alternative to manual annotations for scalable GUI agent development.
发表机构
- Zhejiang University(浙江大学)
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。