从可运行到可验证:LLM/智能体驱动的漏洞验证制品的独立可复现性研究
From Runnable to Verifiable: An Independent Reproducibility Study of LLM/Agent-Driven Vulnerability Validation Artifacts
AI总结:
本研究通过预注册可复现性审计,发现LLM/智能体驱动的漏洞验证制品存在标识符不一致、可运行率低、预言机不可靠等问题,提出的审计方案可作为安全可复现性社区的复用模板。
AI中文摘要:
安全研究制品(代码仓库、概念验证漏洞利用程序、验证流水线)越来越多地由LLM/智能体驱动的漏洞工作流生成,但“公开可用”“可运行”“产生有效信号”“语义确认”的制品之间的差距却缺乏充分衡量。我们对该领域开展了一项预注册的可复现性审计:覆盖2023—2026年的检索经双重筛选后得到104篇论文的共识语料库,其中59篇(56.7%)拥有公开可访问的制品。我们在R0/R1阶段执行了18篇论文的样本,以及锚定基准(arXiv:2509.24037)的全部102个案例,对30个产生信号的案例给出带补丁的反事实裁决,对19个案例给出匹配的阴性对照裁决。研究得出三项关键发现:第一,102个锚定案例中有58个(56.9%)包含脚本内部的CVE标识符,与声明的目录CVE不一致;第二,18篇论文级制品中仅10个(55.6%)能在R0阶段完成声明的工作流,仅修复环境后在R1阶段提升至11个(61.1%);第三,制品内置的预言机不可靠:30项带补丁的反事实审计中,20个在已打补丁的构建版本上仍产生声称的信号,19项匹配阴性对照中,7个在良性输入上仍触发,预言机的混淆矩阵灵敏度为60%,特异度为45%。若没有干净的带补丁反事实对照,在易受攻击的构建版本上触发并不能证明CVE特异性的复现。这些结果来自预注册方案,我们的方案——预注册后置条件、R0/R1修复阶梯、G1—G3语义证据级别、带补丁的反事实预言机——是可复用的安全可复现性社区模板。
英文摘要:
Security research artifacts---repositories, PoC exploits, and validation pipelines---are increasingly produced by LLM/agent-driven vulnerability workflows, yet the gap between \emph{publicly available}, \emph{runnable}, \emph{signal-producing}, and \emph{semantically confirmed} artifacts is poorly measured. We conduct a pre-registered reproducibility audit of this literature. A search covering 2023--2026 with dual screening yields a 104-paper consensus corpus, of which 59 papers (56.7\%) have a publicly reachable artifact. We execute an 18-paper sample at R0/R1 and all 102 cases of the anchor benchmark (arXiv:2509.24037), with patched-counterfactual verdicts on 30 signal-producing cases and matched-negative-control verdicts on 19. Three findings stand out. First, 58/102 (56.9\%) anchor cases contain a script-internal CVE identifier that diverges from the declared directory CVE. Second, only 10/18 (55.6\%) paper-level artifacts complete their declared workflow at R0, rising to 11/18 (61.1\%) after environment-only R1 repair. Third, artifact-embedded oracles prove unreliable: 20/30 patched-counterfactual audits still produce the claimed signal on the patched build, 7/19 matched negative controls still trigger on benign input, and the oracle confusion matrix has sensitivity 60\% and specificity 45\%. A trigger on the vulnerable build is not evidence of CVE-specific reproduction without a clean patched counterfactual. These are exploratory results from a pre-registered protocol, and our protocol---pre-registered post-conditions, R0/R1 repair ladder, G1--G3 semantic evidence levels, and patched-counterfactual oracles---is a reusable template for the security reproducibility community.