arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SeerGuard:一种通过世界模型预测实现的移动 GUI 代理安全框架

SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction

Xue Yu, Bo Yuan, Kailin Zhao, Pengshuai Yang, Hong Hu, Junlan Feng

arXiv 2607.15550首次发表:更新:

发表机构

JIUTIAN Research(中移九天)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对移动 GUI 代理安全风险问题,提出 SeerGuard 框架,通过执行前指令级筛选和行动级风险评估减轻风险。构建安全增强世界模型,经实验验证其在不同代理上有效泛化,提升了安全效用分数并降低风险成本分数。

AI 中文摘要

移动图形用户界面(GUI)代理在自动化复杂任务方面展现出卓越能力,但也带来关键安全风险,单一错误操作可能导致不可逆转的后果。现有安全机制主要是被动反应式的,缺乏执行前评估风险的能力。本文介绍了 SeerGuard,这是一个后果感知安全框架,旨在通过执行前指令级筛选和行动级风险评估来减轻这些风险。具体而言,行动级评估在当前 GUI 状态下分析代理提出的行动,预测可能结果以在执行前识别风险。为实现这些能力,我们通过多任务学习构建了一个统一的安全增强世界模型(SAWM),将语义下一状态预测与安全风险评估相结合。大量实验表明,SeerGuard 能在不同移动 GUI 代理上有效泛化。在 Qwen3 - VL - 8B - Instruct 上,在 ω = 0.8 时安全效用分数从 0.191 提高到 0.596,在 α = 0.8 时风险成本分数从 0.347 降低到 0.130。对我们的 SAWM 的进一步分析验证了指令级筛选的有效性以及行动风险评估和下一状态预测的能力。

英文摘要

Mobile graphical user interface (GUI) agents have demonstrated remarkable capabilities in automating complex tasks, yet they introduce critical safety risks because a single erroneous action can lead to irreversible consequences. Existing safety mechanisms are primarily reactive, lacking the ability to assess risks before execution. In this paper, we introduce SeerGuard, a consequence-aware safety framework designed to mitigate these risks through pre-execution instruction-level screening and action-level risk assessment. Specifically, the action-level assessment analyzes agent-proposed actions within current GUI states, anticipating likely outcomes to identify risks before they are executed. To enable these capabilities, we construct a unified safety-augmented world model (SAWM) via multi-task learning, integrating semantic next-state prediction with safety risk assessment. Extensive experiments demonstrate that SeerGuard generalizes effectively across diverse mobile GUI agents. On Qwen3-VL-8B-Instruct, it increases the safety-utility score from $0.191$ to $0.596$ at $ω=0.8$ and reduces the risk-cost score from $0.347$ to $0.135$ at $α=0.8$. Further analyses on our SAWM validate the effectiveness of the instruction-level screening, alongside the capability of action risk assessment and next-state prediction.

Comments19 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑