arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39547cs.LG

在不完美先验条件下学习可靠的GUI智能体

Learning Reliable GUI Agents under Imperfect Priors

  • MiLM Plus, Xiaomi Inc.(小米公司MiLM Plus)

机构由 AI 辅助整理,请以论文原文为准。

Bo Han, Qianyi Wang, Shuai Liu, Xiong Zifan, Changqiao Wu, Yuanfa Li, Pengzhi Gao, Wei Liu, Jian Luan, Heng Qu, Yunpeng Song, Zhongmin Cai

AI总结:

针对GUI智能体在真实环境中的不完美先验问题,提出结合结构化探索与噪声感知训练框架,在物理设备和模拟器上显著提升任务覆盖率和先验利用准确性。

AI中文摘要:

基于大型语言模型和视觉-语言模型的GUI智能体在未见过的应用和复杂的多步骤任务上仍然表现不佳,因为完成真实的GUI任务依赖于特定应用、随时间变化的操作知识,而这些知识在预训练语料库中稀缺。检索增强执行提供了一种自然的补救方法,但面临两个相互关联的瓶颈:大规模知识的获取困难,以及由于版本更新、促销、广告、A/B测试和个性化,自收集的先验不可避免地偏离实时环境。因此,我们认为GUI智能体不应追求完美的知识,而应学会在不完美先验条件下正确行动,并提出了一个将知识获取与噪声鲁棒利用相结合框架:一种结构化探索策略遍历交互元素,构建UI状态转换图,并通过视觉-语言模型合成(任务,轨迹)对,无需人工标注;一种基于真实GUI漂移模式分类的噪声感知训练策略,将五种类型的现实错误注入自探索轨迹中,以教会智能体在行动前评估先验可靠性。在物理设备和在线模拟器基准上的实验表明,我们的方法发现了更多独特的屏幕,覆盖了更多基准任务,并更有效地拒绝错误先验同时利用正确先验,其准确性提升可跨数据集迁移。

英文摘要:

GUI agents built on large language and vision-language models still struggle on unseen applications and complex multi-step tasks, as completing real GUI tasks depends on app-specific, temporally volatile operational knowledge that is scarce in pretraining corpora. Retrieval-augmented execution offers a natural remedy but faces two coupled bottlenecks: knowledge at scale is hard to acquire, and self-collected priors inevitably drift from the live environment due to version updates, promotions, ads, A/B tests, and personalization. We therefore argue that GUI agents should not pursue perfect knowledge but learn to act correctly under imperfect priors, and propose our framework that couples knowledge acquisition with noise-robust utilization: a structured exploration strategy traverses interactive elements, builds a UI state-transition graph, and synthesizes (task, trajectory) pairs via a VLM without human annotation; a noise-aware training strategy, grounded in a taxonomy of real GUI drift patterns, injects five types of realistic errors into self-explored trajectories to teach the agent to assess prior reliability before acting. Experiments on physical devices and online emulator benchmarks show that our method discovers more unique screens, covers more benchmark tasks, and more effectively rejects erroneous priors while leveraging correct ones, with accuracy gains that transfer across datasets.

补充信息

↑