AI 中文总结
本文通过配对诊断发现GUI智能体的提示级对齐是局部现象,一行防护栏可大幅降低单次ASR,但四轮升级链会使其防护效果下降,静态单轮ASR高估了实际部署的鲁棒性。
AI 中文摘要
GUI智能体在普适计算环境中的可信部署,需要其对齐能力能在动态交互和精确威胁条件下维持,而非仅能单轮拒绝明确的有害请求。本文认为,当前移动智能体中占主导的轻量防御手段——提示级对齐是一种局部现象:它仅在通常被评估的狭窄场景(即单轮、明确表达的意图)中可靠生效,且会沿任何真实用户都能遍历的两个轴系统性退化。通过对三个前沿GUI智能体开展配对诊断,采用基于屏幕、用户侧说服的方式且无环境注入,我们发现一行防护栏可实现单次攻击成功率(ASR)大幅降低,最高约40个百分点,且过度拒绝成本接近零。然而,从独立探测转向四轮升级链后,所有模型的防护ASR均上升约20个百分点。相较于中性基线,该增幅反映出Qwen的防护栏大幅失效,而Claude和GPT的增幅主要是与防御正交的动态风险。显著性差距的符号在防护栏下翻转:无防护栏时,隐藏请求并不比明确请求更易成功;有防护栏时则更易成功,表明防御仅在意图被明确提及才会生效。因此,静态单轮ASR会系统性且可预测地高估部署后的鲁棒性。
英文摘要
Trustworthy deployment of GUI agents in ubiquitous computing settings requires alignment that survives dynamic interaction and precise threat conditions, not just single-turn refusal of explicit harmful requests. We argue that prompt-level alignment, the dominant lightweight defense in current mobile agents, is a local phenomenon: it works reliably only in the narrow evaluation slice where it is typically measured, namely single-turn, explicitly-verbalized intent, and degrades systematically along two axes that any real user can traverse. Using a paired diagnostic on three frontier GUI agents, screen-grounded, user-side persuasion, with no environment injection, we show that a one-line guardrail achieves large single-shot ASR reductions, up to roughly 40 points, at near-zero over-refusal cost. Nevertheless, moving from independent probes to four-turn escalation chains raises guarded ASR by approximately 20 points on every model. Relative to the neutral baselines, this increase reflects substantial guardrail erosion for Qwen but a largely defense-orthogonal dynamic risk for Claude and GPT. The sign of the salience gap flips under the guardrail: concealed requests are not systematically more successful than explicit ones without a guardrail, but are more successful with one, indicating that the defense engages primarily when intent is named. Static single-turn ASR therefore overstates deployed robustness by a systematic and predictable margin.