arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向GUI智能体的软件工程及与GUI智能体协同的软件工程

Software Engineering for and with GUI Agent

Shengcheng Yu, Yuchen Ling, Junyang Xing, Quan Zhou, Chunrong Fang, Zhenyu Chen

arXiv 2608.09278首次发表:更新:

发表机构

State Key Laboratory for Novel Software Technology, Nanjing University; Technical University of Munich(南京大学计算机软件新技术国家重点实验室; 慕尼黑工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过分析2018年1月至2026年4月的336篇GUI智能体论文,指出其在恢复机制、安全执行等方面的不足,提出需结合可靠执行与生命周期测试等以构建可部署的GUI智能体系统。

AI 中文摘要

GUI智能体已取得快速发展,衍生出日益增多的框架、基准及应用。然而,这一发展速度已超过该领域的成熟度,GUI智能体在技术上仍较为脆弱,工程化程度不足,且针对持续实际应用的验证不够充分。它们正逐步演变为闭环软件系统,在这类系统中,模型推理与界面感知、执行反馈、恢复机制及人工监督紧密耦合。这一演进要求采用软件工程视角,但现有研究中基本缺失该视角。为填补这一空白,我们回顾了2018年1月至2026年4月期间的336篇GUI智能体论文,通过5个研究问题考察了该领域的研究现状、架构、评估、软件生命周期相关问题及未来机遇。研究结果显示,自2024年以来该领域已急剧扩张,移动和网页场景仍占主导;架构日益采用模块化的感知-推理-执行循环,但恢复机制、人工 escalation(升级)、安全执行及可审计性仍未得到充分发展。这种架构失衡也延伸至评估环节,评估正变得更具交互性,但仍以任务成功为核心,且难以在不同协议间进行比较。更广泛而言,现有研究对基准之外的测试及智能体发布后的维护支持有限,可观测性、隐私工程及系统性人工监督也未得到充分发展。这些结果共同表明,仅提升能力无法确保部署就绪。未来研究应将可靠执行与以生命周期为中心的测试、可复现的评估相结合,还应将权限和隐私控制与成本感知、以人工为中心的治理相整合,这种整合对于构建可靠、可维护、安全且可部署的GUI智能体系统是必要的。

英文摘要

GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced the maturity of the field. GUI agents remain technically brittle, incompletely engineered, and insufficiently validated for sustained real-world use. They are evolving into closed-loop software systems. Within these systems, model reasoning is coupled with interface perception, execution feedback, recovery, and human oversight. This evolution calls for a software engineering perspective that remains largely absent from existing research. We address this gap by reviewing 336 GUI-agent papers from January 2018 to April 2026. Five research questions examine the research landscape, architectures, evaluation, software lifecycle concerns, and future opportunities. Our findings show that the field has expanded sharply since 2024, while mobile and web settings remain dominant. Architectures increasingly adopt modular perceive-reason-act loops, but recovery, human escalation, safety enforcement, and auditability remain underdeveloped. This architectural imbalance extends to evaluation. Evaluations are becoming more interactive, but they remain centered on task success and are difficult to compare across protocols. More broadly, existing studies provide limited support for testing beyond benchmarks and for maintaining agents after release. Observability, privacy engineering, and systematic human oversight are also underdeveloped. Together, these findings show that capability improvements alone cannot ensure deployment readiness. Future research should connect dependable execution with lifecycle-centered testing and reproducible evaluation. It should also integrate permission and privacy controls with cost-aware, human-centered governance. This integration is necessary to build dependable, maintainable, secure, and deployable GUI-agent systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑