AI 中文总结
研究移动GUI智能体的权限素养,发现其存在应用信任偏差、任务优先级覆盖等问题,提示干预效果不一,建议将任务执行与权限授权分离。
AI 中文摘要
移动GUI智能体在任务执行过程中经常会遇到系统权限对话框,但其仅授予委托任务所需权限的能力在很大程度上未被研究。我们对这一能力开展了系统性研究,将其命名为权限素养。我们基于任务相关性和隐私风险构建了一个四级权限框架,并由三位GUI智能体安全领域的独立专家对评估场景进行验证。我们将Android风格的权限弹窗注入真实GUI任务中,使用同步标注的截图和UI树层次结构评估四个前沿多模态大语言模型,使智能体可获取请求者、权限、正当理由及可用操作等信息。除主要研究外,我们开展了受控干预,分别改变任务上下文和智能体可见的请求者身份。在同一日历(Calendar)任务下,仅将请求者从Calendar更改为PiMusic,授权数量就从26/32降至0/32,这揭示了强烈但受任务制约的应用信任偏差。在弹窗固定的情况下改变任务上下文也会显著改变授权决策,这揭示了系统性的任务优先级覆盖。提示干预可减少不必要的授权,但其效果在不同模型间不一致,且可能以抑制合法授权为代价。这些结果表明,将任务执行与权限授权分离是未来工作的一个有前景的设计方向。
英文摘要
Mobile GUI agents routinely encounter system permission dialogs during task execution, yet their ability to grant only permissions that are necessary for the delegated task remains largely unexamined. We present a systematic study of this capability, which we term Permission Literacy. We construct a four-level permission framework based on task relevance and privacy risk and validate the evaluated scenarios with three independent experts in GUI-agent safety. We inject Android-style permission popups into real GUI tasks and evaluate four frontier multimodal large language models using synchronized annotated screenshots and UI-tree hierarchies, making the requester, permission, justification, and available actions accessible to the agent. Beyond the main study, we conduct controlled interventions that separately vary task context and agent-visible requester identity. Under the same Calendar task, changing only the requester from Calendar to PiMusic reduces grants from 26/32 to 0/32, revealing a strong but task-conditioned App-Trust Bias. Holding a popup fixed while changing task context also substantially changes authorization decisions, revealing a systematic Task-Prior Override. Prompt interventions can reduce unnecessary grants, but their effectiveness is inconsistent across models and may come at the cost of suppressing legitimate grants. These results suggest that separating task execution from permission authorization is a promising design direction for future work.