arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27030cs.CRcs.LG

HoF-Bench:无需前沿模型即可重新发现真实AI发现的CVE

HoF-Bench: Rediscovering Real AI-Discovered CVEs Without Frontier Models

Petr Simecek, Elnaz Babayeva, Jiri Balhar, Michal Bida, Michal Buran, Vaclav Cadek, Luigino Camastra, Tomas Dulka, Michal Janocko, Tomas Klohna, Pavel Kohout, O… 展开作者

Petr Simecek, Elnaz Babayeva, Jiri Balhar, Michal Bida, Michal Buran, Vaclav Cadek, Luigino Camastra, Tomas Dulka, Michal Janocko, Tomas Klohna, Pavel Kohout, Ondrej Kokes, Adam Krivka, Jakub Kubik, Patrik Mada, Igor Morgenstern, Marek Pavelka, Joshua Rogers, Petr Stastny, Jan Tattermusch, Dmitrijs Trizna, Martin Votruba, Guido Vranken, Jakub Zikl, Evelina Gabasova, Stanislav Fort

首次发表
浏览论文内容

中文总结 AI 辅助

HoF-Bench基准基于AISLE公开的95个AI发现的CVE构建,可用于对比不同模型的漏洞检测能力,极简LLM分析器在严格协议下可重新发现68%的CVE,所有模型未发现的CVE集中在C基础设施代码中。

中文摘要 AI 辅助

基于大语言模型(LLM)的分析器已开始在成熟开源项目中发现真实漏洞:AISLE的分析器在OpenSSL、curl、GnuTLS等78个项目中发现了280多个CVE。我们推出HoF-Bench(命名自AISLE的公开名人堂),该基准由8个代码库中95个此类公开AI发现的CVE构建,且这些代码库均固定在易受攻击的提交版本上。分析器仅接收源文件和目标文件范围,不提供CVE标识符、描述、修复方案或预期机制;由检测器盲态的前沿模型评判者仅认可那些能识别相同代码路径、根本原因、攻击条件和影响的发现。在该严格协议下,一个刻意设计的极简LLM分析器最多可重新发现95个CVE中的65个(占比68%)。本研究中没有前沿模型执行检测任务。10种检测器骨干包括5个开放权重模型(总参数规模21B至284B,激活参数3B至13B)和5个专有小型或“闪速”级模型,所有模型均在固定框架中运行,包含4次重复传递、可选的生成上下文阶段以及可重复的多轮分类阶段(共7600条模型-CVE传递记录)。难度受语言影响显著,所有模型均未发现的CVE集中在C基础设施代码中。HoF-Bench为比较漏洞扫描器、其重复运行的可靠性以及它们产生的候选量提供了紧凑的测试平台,该数据集可在此httpsURL获取。

英文摘要

LLM-based analyzers have begun finding real vulnerabilities in mature open-source projects: AISLE's analyzer is credited with more than 280 CVEs across 78 projects, including OpenSSL, curl, and GnuTLS. We introduce HoF-Bench (named after AISLE's public Hall of Fame), a benchmark built from 95 of these public AI-discovered CVEs across eight repositories pinned at vulnerable commits. Analyzers receive source and target-file scope but not CVE identifiers, descriptions, fixes, or expected mechanisms; a detector-blinded frontier-model judge credits only findings that identify the same code path, root cause, attack condition, and impact. A deliberately minimal LLM-based analyzer rediscovers up to 65 of the 95 CVEs (68%) under this strict protocol. No frontier model performs detection anywhere in the study. The ten detector backbones are five open-weight models (21B--284B total parameters, 3--13B active) and five proprietary small or "flash"-tier models. All of them run in the fixed scaffold with four repeated passes, an optional generated-context stage, and a replayable multi-round triage stage (7,600 model--CVE pass records). Difficulty is strongly structured by language; the CVEs missed by every model concentrate in C infrastructure code. HoF-Bench provides a compact test bed for comparing vulnerability scanners, their reliability across repeated runs, and the candidate volume they create. The dataset is available at https://huggingface.co/datasets/aisleinc/HoF-Bench.

发表机构

  • AISLE
  • Central European Institute of Technology(中欧技术研究所)
  • Masaryk University(马萨里克大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑