arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高效GUI智能体:关于观察、记忆、动作与运行时优化的系统综述

Efficient GUI Agents: A Systems Survey of Observation, Memory, Action, and Runtime Optimization

Bizhe Bai, Jiakang Yuan, Hongming Wu, Xinyue Wang, Jie Ren, Siyao Chen, Yuchen Ya, Fan Bai, Pai Peng, Huafeng Qin, Tao Chen

arXiv 2609.02309首次发表:更新:

发表机构

College of Future Information Technology, Fudan University; Shanghai Innovation Institute; Chongqing Technology and Business University(复旦大学未来信息技术学院; 上海创新研究院; 重庆工商大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本综述从系统视角梳理高效GUI智能体的技术维度,总结其核心进展思路,并指出验证器成本核算等开放问题,为该领域实际部署提供参考。

AI 中文摘要

GUI智能体越来越多地在网站、移动应用和桌面环境中运行,但该领域的进展主要仍通过任务成功率来体现。我们认为,实际部署同样取决于效率:智能体在完成任务时消耗的上下文、计算量、动作预算和运行时开销。本综述从端到端的系统视角研究高效GUI智能体,涵盖当前技术维度:观察效率、上下文与记忆效率、动作效率以及规划器侧/系统效率。对于每个子部分,我们通过定向搜索及前后引用链扩展核心文献,进而综合主流机制、报告的效率信号及它们引入的新开销。现有文献中,近期进展汇聚于少量反复出现的思路:选择性读取而非全上下文摄取、全局到局部的视觉分配、可恢复记忆而非原始历史回放、验证感知控制,以及可在GUI与非GUI执行间切换的混合运行时。我们最后指出主要开放问题,包括验证器成本的如实核算、跨基准可比性,以及在实际延迟与隐私约束下观察、记忆和执行层的协同设计。

英文摘要

GUI agents increasingly operate across websites, mobile apps, and desktop environments, yet the field still reports progress primarily through task success. We argue that practical deployment depends equally on efficiency: how much context, computation, action budget, and runtime overhead an agent consumes while succeeding. This survey studies efficient GUI agents through an end-to-end systems lens that preserves the current technical axes of observation efficiency, context and memory efficiency, action efficiency, and planner-side/system efficiency. For each subsection, we expand the seed literature through targeted search plus backward and forward citation chaining, then synthesize the dominant mechanisms, reported efficiency signals, and new overheads they introduce. Across the literature, recent progress converges on a small set of recurring ideas: selective reading instead of full-context ingestion, global-to-local visual allocation, recoverable memory rather than raw history replay, verification-aware control, and hybrid runtimes that can switch between GUI and non-GUI execution. We conclude by identifying the main open problems, including honest accounting of verifier cost, cross-benchmark comparability, and co-design of observation, memory, and execution layers under real latency and privacy constraints.

CommentsAccept at Grounding Language Models: Learning Faithfully and Efficiently @ EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑