ReproBench:基准测试LLM智能体从零开始复现漏洞
ReproBench: Benchmarking LLM Agents on Reproducing Vulnerability From Scratch
- Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
- University of Chinese Academy of Sciences(中国科学院大学)
- Key Laboratory of System Software (Chinese Academy of Sciences)(中国科学院系统软件重点实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出ReproBench基准,从CVE标识符出发评估LLM智能体端到端漏洞复现能力,发现仅5.3%成功,但证实了自主复现的潜力,并识别了关键瓶颈。
AI中文摘要:
大型语言模型(LLM)智能体越来越多地被评估于网络安全任务,如漏洞复现、利用和修补。然而,现有的网络安全基准主要在环境后评估范式下运行,即向智能体提供源代码、容器或可执行二进制文件。这种设置绕过了关键的环境重建步骤,留下了一个真实世界漏洞分析中的基本问题:智能体能否自主重建所需的执行环境并完全从零开始复现漏洞?为解决这一空白,我们提出了ReproBench,一个基于证据的基准,旨在评估智能体在仅从CVE标识符出发进行端到端漏洞复现的能力。ReproBench将完整的复现工作流分解为六个不同阶段,并使用可验证的实验产物(如下载的固件镜像、解包的二进制文件、细粒度的分析日志和验证的崩溃样本)独立评估每个阶段的性能。我们用30个真实世界的物联网固件漏洞实例化ReproBench,这些漏洞作为我们从零开始评估设置的理想测试用例。我们的评估表明,45.3%的测试运行诉诸于漏洞模拟——一种在所有被评估的LLM智能体中普遍采用的修复变通方法——而只有5.3%的CVE-模型对成功复现了真实世界的漏洞。尽管总体成功率较低,但这些非平凡的成功案例证实了最先进的LLM智能体已经具备完全自主的端到端漏洞复现能力。同时,我们对失败案例的深入分析识别出阻碍LLM智能体在复现流程中的核心瓶颈,为后续研究提供了可操作的见解。
英文摘要:
Large language model (LLM) agents are increasingly evaluated on cybersecurity tasks such as vulnerability reproduction, exploitation, and patching. However, existing cybersecurity benchmarks predominantly operate under a post-environment evaluation paradigm, i.e., handing the agent source code, a container, or an executable binary. This setup bypasses the critical environment reconstruction step, leaving a fundamental question for real-world vulnerability analysis: can an agent autonomously reconstruct the required execution environment and reproduce a vulnerability entirely from scratch? To address this gap, we present ReproBench, an evidence-grounded benchmark designed to evaluate agent capabilities in end-to-end vulnerability reproduction starting from solely a CVE identifier. ReproBench decomposes the full reproduction workflow into six distinct phases, and assesses performance on each phase independently using verifiable experimental artifacts: downloaded firmware images, unpacked binaries, granular analysis logs, and validated crash samples, among others. We instantiate ReproBench with 30 real-world IoT firmware vulnerabilities, which serve as ideal test cases for our from-scratch evaluation setting. Our evaluation demonstrates that 45.3% of test runs resort to vulnerability simulation - a prevalent remediation workaround adopted across all evaluated LLM agents - while only 5.3% of CVE-model pairs achieve successful reproduction of real-world vulnerabilities. Despite the low overall success rate, these non-trivial successful cases confirm that state-of-the-art LLM agents already possess the capacity for fully autonomous end-to-end vulnerability reproduction. Concurrently, our in-depth analysis of failed cases identifies core bottlenecks impeding LLM agents throughout the reproduction pipeline, offering actionable insights for subsequent research.