发表机构
The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DIBench提出决策完整性基准,通过欺骗性注入评估移动智能体在多候选选择中的目标偏离风险,实验显示完成率评估会掩盖安全隐患。
AI 中文摘要
随着基于图形用户界面(GUI)的移动智能体快速发展,对其在真实应用界面中自主决策进行严格的安全评估变得日益关键。现有基准主要关注执行层面的异常,使用任务成功率或劫持率等指标,但未能捕捉多候选选择任务中的任务内目标偏离风险,在这种任务中,决策可能被引导至攻击者指定的目标,甚至违反指令隐含的约束(如最便宜/评分最高),且不伴随任何明显的执行异常。我们提出DIBench,一个用于衡量移动智能体中此类风险的决策完整性基准。DIBench涵盖7个商业应用和3个模拟应用,包含5种任务类型。在仅限于非特权UI内容的威胁模型下,我们构建了8种欺骗性注入探针实例,这些实例能够在无明显异常的情况下引导关键选择。该基准包含1,000个干净实例和36,672个注入实例,并提供统一的协议和完整性指标以供比较。涵盖4个智能体框架和7个基础模型的实验表明,基于完成度的评估可能高估智能体的可信度,并遗漏决策完整性风险:欺骗性注入会引导选择并改变早期行动策略,从而虚增完成率并造成误导性的安全假象。常见的防御措施,包括检测、图像预处理和提示提醒,在完整性提升方面效果不一致。总体而言,DIBench提供了一个统一、可复现的基准,用于量化移动智能体中任务内目标偏离的风险,并支持对安全防御措施进行可比较的评估。
英文摘要
As GUI-based mobile agents rapidly progress, rigorous safety evaluation of their autonomous decision-making in realistic app interfaces becomes increasingly critical. Existing benchmarks mainly focus on execution-level anomalies using task success or hijack rates, but fail to capture the in-task goal deviation risk in multi-candidate selection tasks, where the decision may be steered toward an attacker-specified target, even in violation of instruction-implied constraints (e.g., cheapest/highest-rated), without any overt execution anomalies. We present DIBench, a decision integrity benchmark for measuring this risk in mobile agents. DIBench covers 7 commercial and 3 simulated apps with 5 task types. Under a threat model restricted to non-privileged UI content, we construct 8 deceptive injection probe instantiations that can steer critical selections without overt anomalies. The benchmark includes 1,000 clean and 36,672 injected instances, with a unified protocol and integrity metrics for comparison. Experiments spanning 4 agent frameworks and 7 base models show that completion-based evaluation can overestimate agent trustworthiness and miss decision-integrity risks: deceptive injections steer selections and shift early action policies, inflating completion rates and creating a misleading illusion of safety. Common defenses, including detection, image preprocessing, and prompt reminders, yield inconsistent integrity gains. Overall, DIBench provides a unified, reproducible benchmark to quantify the risk of in-task goal deviation in mobile agents and enable comparable evaluations of safety defenses.
CommentsAccepted to NeurIPS 2026, Track on Evaluations and Datasets