发表机构
University of Illinois Urbana-Champaign; Ministry of National Defense, Republic of Korea(伊利诺伊大学厄巴纳-香槟分校; 韩国国防部)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出KEX-bench基准,评估编码智能体生成内核利用原语的能力,结果显示无参考PoC时成功率低,凸显从崩溃到利用的差距。
AI 中文摘要
编码智能体如今能够在生产软件中发现真实漏洞。然而,漏洞发现结果并不能衡量智能体能否构建利用原语。我们引入了KEX-bench,一个用于评估编码智能体在针对真实操作系统内核的利用原语生成方面表现的基准。KEX-bench包含45个任务实例,覆盖40个Linux和Windows CVE,涉及内核地址泄露、指令指针控制、堆读取、堆写入和任意地址写入。每个任务在隔离的虚拟机中运行,暴露受控工具,并使用确定性验证器检查特定原语的成功。我们在固定工具调用预算下评估了配备前沿和开放权重模型的最先进编码智能体。在没有参考概念验证(PoC)的情况下,最强配置解决了20个Windows任务中的1个(5.0%)和25个Linux任务中的14个(56.0%)。在提供参考PoC的情况下,最强配置解决了45个任务中的31个(68.9%)。这凸显了智能体能够达到内核崩溃但未能将内核状态塑造成利用原语的差距。我们发布KEX-bench,以促进在https URL上对AI辅助利用的可复现研究。
英文摘要
Coding agents now find real vulnerabilities in production software. However, bug discovery results do not measure whether agents can construct exploit primitives. We introduce KEX-bench, a benchmark for evaluating coding agents on exploit primitive generation against real operating-system kernels. KEX-bench contains 45 task instances across 40 Linux and Windows CVEs, covering kernel address leak, instruction-pointer control, heap read, heap write, and arbitrary address write. Each task runs in an isolated virtual machine, exposes controlled tools, and uses a deterministic verifier to check primitive-specific success. We evaluate state-of-the-art coding agents paired with frontier and open-weight models under fixed tool-call budgets. Without a reference proof of concept (PoC), the strongest configuration solves 1 of 20 Windows tasks (5.0%) and 14 of 25 Linux tasks (56.0%). With a reference PoC, the strongest configuration solves 31 of 45 tasks (68.9%). This highlights the gap where agents reach kernel crashes but fail to shape kernel state into exploit primitives. We release KEX-bench for reproducible research on AI-assisted exploitation at https://kex-bench.github.io.