发表机构
The Hong Kong Polytechnic University; Huazhong University of Science and Technology(香港理工大学; 华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示移动智能体与人类用户因感知差异导致的UI失同步威胁,提出自动化攻击框架,在五个框架上实现高达77.9%的误导率,且扰动难以被人类察觉。
AI 中文摘要
移动智能体日益能够自主地与移动应用交互,并代表用户执行具有重要后果的操作。对此类智能体的有效人类监督依赖于一个基本前提:用户和智能体从同一界面观察到一致的信息。我们证明,这一前提可能被系统性破坏。用户通过物理显示屏和人类视觉系统感知移动界面,因此其观察会受到遮挡和亮度对比度限制的影响。相比之下,智能体消费的数字截图可能保留此类内容,而可访问性表示则暴露非视觉的小部件元数据。因此,相同的UI状态可能向用户和智能体呈现实质不同的信息,我们将这种不匹配称为人机UI失同步。我们研究了一个经过重新打包的合法APK克隆是否可以利用这种失同步将智能体引导至攻击者指定的操作,同时对于人类用户而言,该克隆保持完全功能且与原应用行为一致。我们证明这一威胁是可行的:部署前嵌入的扰动可以在无需访问运行时用户指令、智能体检测或在线自适应的情况下引发此类偏差。为了系统地暴露和评估这一威胁,我们开发了一个自动化框架,该框架构建与用户运行时指令无关的UI失同步攻击,并在可部署的APK中实现它们。我们在五个移动智能体框架和三个骨干模型上,对涉及各种应用的546项任务进行了静态和动态评估,分别实现了77.9%和66.9%的平均误导率。一项包含186名参与者的补充问卷研究发现,我们攻击中使用的视觉扰动对人类用户而言难以察觉。
英文摘要
Mobile agents are increasingly capable of autonomously interacting with mobile applications and performing consequential actions on behalf of users. Effective human oversight of such agents relies on a basic premise: users and agents observe consistent information from the same interface. We show that this premise can be systematically violated. Users perceive mobile interfaces through physical displays and the human visual system, making their observations subject to occlusion and luminance contrast limitations. In contrast, agents consume digital screenshots that may retain such content and accessibility representations that expose nonvisual widget metadata. The same UI state can therefore present materially different information to users and agents, a mismatch we term human-agent UI desynchronization. We investigate whether a repackaged clone of a legitimate APK can exploit this desynchronization to steer an agent toward attacker-designated actions, while remaining fully functional and behaviorally consistent with the original application for human users. We demonstrate that this threat is feasible: perturbations embedded before deployment can induce such deviations without access to runtime user instructions, agent detection or online adaptation. To systematically expose and evaluate this threat, we develop an automated framework that constructs user runtime instruction-agnostic UI desynchronization attacks and realizes them in deployable APKs. We conduct static and dynamic evaluations across five mobile-agent frameworks and three backbone models on 546 tasks involving various applications, achieving average misleading rates of 77.9% and 66.9%, respectively. A complementary questionnaire-based study with 186 participants finds that the visual perturbations used in our attacks are difficult for human users to notice.
Comments18 pages, 9 figures