arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15457cs.CV

AnchorGUI:GUI导航中双尺度学习的不对称记忆

AnchorGUI: Asymmetric Memory for Dual-Scale Learning in GUI Navigation

Shengjie Jin, Zelong Sun, Hengbo Xu, Yanbiao Ma, Zhiwu Lu

首次发表
浏览论文内容

中文总结 AI 辅助

AnchorGUI利用认知状态锚点生成预测误差信号,通过不对称记忆实现试验内纠正与跨试验提炼,在AndroidWorld上取得57.3%成功率并降低2.4倍令牌消耗。

中文摘要 AI 辅助

视觉语言模型(VLMs)能够实现自主GUI导航,但智能体在处理和学习密集、连续的视觉历史方面仍面临困难。这一瓶颈阻碍了单次尝试(试验内)中的即时错误纠正以及多次尝试(跨试验)中的经验提炼。我们将这些挑战归因于GUI导航中经验性的信息不对称:虽然预期的转换通常可以压缩为轻量级的文本摘要,但意外的结果则受益于保留截图作为准确诊断的因果证据。基于这一见解,我们提出了AnchorGUI,一个由认知状态锚点(CSA)驱动的统一框架。CSA充当每步的原始单元,主动比较预期与观察到的转换,将被动的多模态轨迹转化为显式的预测误差信号。这些信号通过不对称记忆协调双尺度学习机制。对于试验内纠正,滑动窗口选择性地保留检测到不匹配的视觉证据,提供即时的、基于视觉的反馈。对于跨试验提炼,这种不对称记忆将计算昂贵的信用分配搜索空间聚焦于可能的失败步骤。在四个基准上的实验验证了我们方法的有效性。在AndroidWorld上,AnchorGUI实现了57.3%的成功率,每步令牌减少2.4倍。此外,跨试验提炼达到了69.2%的成功率(+11.9%的提升),显著优于标准反思方法,同时保持了次线性的上下文扩展。

英文摘要

Vision-Language Models (VLMs) enable autonomous GUI navigation, but agents still struggle to process and learn from dense, continuous visual histories. This bottleneck hinders both immediate error correction within a single episode (intra-trial) and experience distillation across multiple attempts (cross-trial). We trace these challenges to an empirical informational asymmetry in GUI navigation: while expected transitions can often be compressed into lightweight textual summaries, unexpected outcomes benefit from preserved screenshots as causal evidence for accurate diagnosis. Building on this insight, we propose AnchorGUI, a unified framework driven by the Cognitive State Anchor (CSA). The CSA acts as a per-step primitive that actively compares expected and observed transitions, converting passive multimodal trajectories into explicit prediction-error signals. These signals orchestrate a dual-scale learning mechanism via an asymmetric memory. For intra-trial correction, a sliding window selectively retains visual evidence for detected mismatches, providing immediate, visually-grounded feedback. For cross-trial distillation, this asymmetric memory focuses the computationally expensive credit assignment search space on likely failure steps. Experiments across four benchmarks validate the effectiveness of our approach. On AndroidWorld, AnchorGUI achieves a 57.3% success rate with a $2.4\times$ token reduction per step. Furthermore, cross-trial distillation reaches 69.2% success (+11.9% gain), significantly outperforming standard reflection methods while maintaining sub-linear context scaling.

发表机构

  • Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学高瓴人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑