发表机构
Stanford University; UC Berkeley(斯坦福大学; 加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出MobileCybench基准,用可执行探针评估AI智能体在Android应用中的漏洞发现能力,实验表明顶级智能体在恶意应用场景下触发率53.8%,并发现23个新漏洞。
AI 中文摘要
AI智能体现在报告漏洞的速度比维护者审查的速度还要快。报告通常依赖于应用程序特定的安全属性,并且需要大量人工来处理。为了缓解这一问题,我们引入了一个通过探针(即安全属性的可执行检查)来评估漏洞报告的框架。一份报告的漏洞利用会通过在应用程序上重放并运行探针来评估:触发探针既表明漏洞利用成功,也表明它违反了哪个安全属性。由于探针编码的是安全属性而非已知漏洞,因此它可以检测出编写探针时未知的漏洞。我们将该框架实例化为MobileCybench,这是一个针对13个Android应用中AI智能体漏洞发现的基准测试,包含495个由作者编写和审查的探针。我们在4种设置下评估了5个编码智能体(OpenCode with GPT-5.5、GPT-5.6-Sol和GLM-5.2;Claude Code with Opus 4.8和Opus 5):作为受害者设备上的恶意应用,或作为拥有低权限账户的远程攻击者,每种情况要么只有混淆的APK,要么可以访问应用程序的源代码。仅给定混淆的APK,顶级智能体OpenCode with GPT-5.6-Sol在恶意应用设置中触发了53.8%应用的探针,在远程攻击者设置中触发了16.7%应用的探针。有了源代码,所有智能体和两种攻击设置的触发率从28.8%增加到32.8%。构建和运行该基准测试发现了23个先前未报告的漏洞,其中大多数已得到维护者的确认。
英文摘要
AI agents now report vulnerabilities faster than maintainers can review them. Reports often depend on security properties specific to the application, and require considerable human labor to process. To mitigate this, we introduce a framework for evaluating vulnerability reports via probes, executable checks of security properties. A reported exploit is evaluated by replaying it against the application and running the probes: a triggered probe indicates both that the exploit succeeded and which security property it violated. As a probe encodes a security property rather than a known vulnerability, it can detect vulnerabilities that were not known when the probe was written. We instantiate the framework as MobileCybench, a benchmark for vulnerability discovery by AI agents in 13 Android applications, with 495 probes written and reviewed by the authors. We evaluate 5 coding agents (OpenCode with GPT-5.5, GPT-5.6-Sol, and GLM-5.2; Claude Code with Opus 4.8 and Opus 5) under 4 settings: as a malicious app on the victim's device or as a remote attacker with a low-privilege account, each with either only an obfuscated APK or access to the application's source code. Given only the obfuscated APK, the top agent, OpenCode with GPT-5.6-Sol, triggers probes in 53.8% of applications in the malicious-app setting and 16.7% in the remote-attacker setting. With source code, the trigger rate across all agents and both attack settings increases from 28.8% to 32.8%. Building and running the benchmark surfaced 23 previously unreported vulnerabilities, the majority of which have been confirmed by maintainers.