AI 是否帮助网络攻击者或防御者?来自非公开漏洞及后续攻击的证据
Does AI Help Cyber Attackers or Defenders? Evidence from Nonpublic Vulnerabilities and Subsequent Attacks
浏览论文内容
中文总结 AI 辅助
本研究通过非公开漏洞环境评估 AI 模型在漏洞利用与修复上的能力,发现不同系统表现差异显著,且初始防御成功不足以阻止后续攻击,强调需进行漏洞特定的攻击-修复比较与后续抵抗测试。
中文摘要 AI 辅助
前沿人工智能系统的发布决策日益依赖于网络能力基准,然而公开的漏洞基准可能使智能体接触到先前已发布的公告、漏洞利用代码和修复补丁,从而难以区分先前暴露与对未见漏洞的实际能力。我们在五个非公开软件环境中评估了开放权重和专有 AI 模型在漏洞利用生成、漏洞修复及后续攻击方面的表现,其中包括我们在漏洞尚未修补时私下披露的漏洞。由研究人员开发并审查的确定性评分器(而非 LLM 评判员)决定任务得分。与 209 个已披露漏洞及密码学挑战的比较揭示了不同系统和漏洞类型之间的显著差异。在两个非公开环境中,修复得分高于攻击得分,而在三个环境中则低于攻击得分。通过初始安全测试也不足以证明安全性:在初始漏洞利用被阻止后,另一个漏洞利用在 524 个非独立防御者测试区间中的 92 个中成功。这些结果促使进行针对特定漏洞的攻击-修复比较及后续抵抗测试。
英文摘要
The release decision for frontier AI systems increasingly relies on cyber capability benchmarks, yet public vulnerability benchmarks can expose agents to previously published advisories, exploits, and fixes, making it difficult to distinguish prior exposure from capability on unseen vulnerabilities. We evaluate open-weight and proprietary AI models on exploit generation, vulnerability repair and subsequent attacks in five nonpublic software environments, including vulnerabilities we privately disclosed while they remained unpatched. Researcher-developed and reviewed deterministic graders, not LLM judges, determine task scores. Comparisons with 209 disclosed vulnerabilities and cryptographic challenges reveal substantial variation across systems and vulnerability types. Repair scores exceed attack scores in two nonpublic environments and fall below them in three. Passing an initial security test is also insufficient: another exploit succeeds in 92 of 524 non-independent defender test intervals after the initial exploit is stopped. These results motivate vulnerability-specific attack-repair comparisons and subsequent resistance tests.
发表机构
- XOR
- Protege
- CISPA Helmholtz Center for Information Security(CISPA亥姆霍兹信息安全中心)
机构由 AI 辅助整理,请以论文原文为准。