arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26314cs.CRcs.AI

StealthBench:衡量自主进攻性安全智能体的操作隐蔽性

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents

Ads Dawson, Adrian Wood

首次发表
浏览论文内容

中文总结 AI 辅助

研究人员推出StealthBench基准,通过含三个LLM的评判小组评估自主进攻性安全智能体的OPSEC表现,发现无模型安全成功率超54%,证实OPSEC失败具系统性,该基准用于支持相关智能体开发与监控。

中文摘要 AI 辅助

隐蔽性指在达成目标时不暴露自身存在、能力或收集到的情报,这是区分专业操作者与可被检测者的关键。顶尖安全研究人员和高级持续威胁(APT)能在不被察觉的情况下达成目标,自主智能体也日益承担相同的进攻任务,但它们是否继承了相关操作技巧?我们推出StealthBench,这是一个衡量自主进攻性安全智能体在六个操作安全(OPSEC)维度上操作隐蔽性的基准。我们从真实的漏洞赏金和红队轨迹中提取了11个经人工验证的OPSEC事件,扩展为14个Docker化任务场景,其中智能体虽发现了真实漏洞,但出现了不符合标准操作技巧的隐蔽性失败:将凭证嵌入公共上传内容、删除生产资源以证明访问权限、强制添加无关用户以展示竞态条件。我们使用由三个模型组成的大语言模型(LLM)评判小组,通过多数投票聚合来评估智能体轨迹,测量安全成功率(完成任务且隐蔽)、Stealth@Solve(成功完成任务中的操作技巧质量)和鲁莽完成率(完成任务但暴露)。结果显示,没有模型的安全成功率超过54%(这一复合指标要求同时满足任务完成和隐蔽),证实OPSEC失败在各模型家族中普遍存在。我们将StealthBench作为公共基准发布,以支持隐蔽感知智能体的开发以及自主进攻性安全部署的自动化OPSEC监控。交互式排行榜、评估工具和数据集可在此httpsURL获取。

英文摘要

Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisticated operators from detectable ones. Elite security researchers and advanced persistent threats achieve their objectives unnoticed; autonomous agents increasingly inherit the same offensive tasks, but do they inherit the tradecraft? We introduce StealthBench,a benchmark that measures operational stealth in autonomous offensive-security agents across six operational security (OPSEC) dimensions. We extract 11 hand-verified OPSEC incidents from real bug-bounty and red-team trajectories, expanded into 14 dockerized task scenarios, where agents, despite finding real vulnerabilities, committed stealth failures inconsistent with standard operational tradecraft: embedding credentials in public uploads, deleting production resources to prove access, force-adding uninvolved users to demonstrate a race condition. We evaluate agent trajectories using a 3-model large language model (LLM) judge panel with majority-vote aggregation, measuring safe success rate (solved and stealthy), Stealth@Solve (tradecraft quality among successful solves), and reckless solve rate (solved but cover blown). Our results show that no model exceeds 54% safe success rate (the compound metric requiring both task completion and stealth), confirming that OPSEC failures are systematic across model families. We release StealthBench as a public benchmark to support both the development of stealth-aware agents and automated OPSEC monitoring for autonomous offensive-security deployments. The interactive leaderboard, evaluation harness, and dataset are available at https://stealthbench.com.

补充信息

↑