发表机构
NYU Tandon School of Engineering; NYU Abu Dhabi; CISPA - Helmholtz Center for Information Security; International Institute of Information Technology Hyderabad; Indian Institute of Technology Tirupati(纽约大学坦登工程学院; 纽约大学阿布扎比分校; CISPA亥姆霍兹信息安全中心; 海得拉巴国际信息技术学院; 蒂鲁帕蒂印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有CTF基准评估智能体进攻性安全能力时忽略轨迹的问题,本文提出CTF-ABACUS框架,通过追踪级审计验证漏洞利用,发现仅62%-87%的恢复Flag为真实利用,为优化基准设计提供依据。
AI 中文摘要
夺旗赛(Capture-the-Flag, CTF)基准被广泛用于评估自主式语言模型智能体的进攻性安全能力。现有评估依赖浅层二元判断或聚合分数,却忽略了智能体获取Flag的轨迹,导致实际漏洞利用与直接Flag暴露、记忆召回、外部查询、猜测及无依据的主张被混为一谈,可能夸大智能体的网络安全能力。本文提出CTF-ABACUS,一种基于追踪的智能体审计框架,它将每次运行重构为基于证据的求解档案。通过将智能体动作分解为渗透测试阶段和分类技术,该框架可识别漏洞利用发生的位置、Flag首次出现的位置,以及恢复的Flag是否有已展示行为的支持。聚合不同智能体的求解档案可生成挑战特征,揭示成功是通过预期漏洞利用还是捷径路径实现的。我们将CTF-ABACUS应用于6个前沿及开源模型在240个挑战上的1435次CTF尝试,在两个评判视角下得到2870份求解档案。追踪验证的漏洞利用仅占基准中恢复Flag的62%-87%,而捷径恢复的轨迹明显更浅。这些发现将CTF评估从计数恢复Flag转变为验证已展示的漏洞利用,并为设计能更好隔离进攻性能力的基准提供了基础。
英文摘要
Capture-the-Flag (CTF) benchmarks are widely used to assess the offensive security capabilities of autonomous language-model agents. Evaluations rely on shallow binary judgments or aggregate scores, overlooking the agent's trajectory to the flag. Consequently actual exploitation is conflated with direct flag exposure, memorized recall, external lookup, guessing, and unsupported claims, potentially overstating the agent's cybersecurity capability. We introduce CTF-ABACUS, a trace-based agent auditing framework that reconstructs each run as an evidence-grounded solve profile. By decomposing agent actions into penetration-testing phases and categorical techniques, it identifies where exploitation occurs, where the flag first appears, and whether the recovered flag is supported by demonstrated behavior. Aggregating solve profiles across agents yields challenge signatures that reveal whether success was achieved via the intended exploit or via shortcut pathways. We apply CTF-ABACUS to 1,435 CTF attempts by six frontier and open-source models on 240 challenges, yielding 2,870 solve profiles under two judge lenses. Trace-verified exploits account for only 62-87% of recovered flags across benchmarks, while shortcut recoveries follow substantially shallower trajectories. These findings shift CTF evaluation from counting recovered flags to verifying demonstrated exploitation and provide a basis for designing benchmarks that better isolate the offensive capabilities.