arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25158cs.AIcs.CRcs.LGcs.SE

FuzzingBrain-Bench V1:评估大语言模型的开放式漏洞发现能力

FuzzingBrain-Bench V1: Evaluating Open-Ended Bug Discovery by LLMs

Ze Sheng, Aleksandar Kezic, Zhicheng Chen, Jeff Huang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出FuzzingBrain-Bench V1基准,评估Claude系列大模型的开源软件漏洞发现能力,其中Claude Opus 4.8表现最佳,在77项挑战中60项触发崩溃,得196分。

中文摘要 AI 辅助

评估大语言模型(LLMs)发现软件漏洞的能力愈发重要。现有基准通常要求模型生成触发预定义目标漏洞的概念验证输入来评估该能力,但这种设置可能会忽略模型发现的与预定义目标不匹配的有效崩溃,导致评估无法反映模型的真实能力。我们提出FuzzingBrain-Bench,这是一个评估AI模型在开源软件中发现漏洞能力的基准。模型将获得一个开源项目和一个在独立Docker镜像中的 sanitizer 检测工具,目标是通过该工具生成输入以触发尽可能多的不同崩溃。模型在每个挑战中的性能根据其产生的不同崩溃签名数量评分,评分上限为预定义最大值并乘以难度系数。FuzzingBrain-Bench V1包含来自43个开源项目的77个挑战,其中36个为C语言、32个为C++语言、9个为Java/JVM语言挑战。我们在整个基准上评估了Claude Haiku 4.5、Claude Sonnet 4.6和Claude Opus 4.8,结果显示Claude Opus 4.8表现最佳,在77个挑战中的60个触发了崩溃,得分196(满分579),三个模型均未在13个挑战中触发崩溃。FuzzingBrain-Bench语料库和检测工具可在该https URL公开获取。

英文摘要

Evaluating the ability of large language models (LLMs) to discover software bugs is increasingly important. Existing benchmarks typically evaluate this capability by asking the model to generate a proof-of-concept input that triggers a predefined target vulnerability. However, this setup may overlook valid crashes discovered by the model when they do not match the predefined target. As a result, the evaluation may not reflect the model's real capability. We present FuzzingBrain-Bench, a benchmark for assessing AI models' ability to discover bugs in open-source software. Models are given an open-source project and a sanitizer-instrumented harness in a self-contained Docker image. Their goal is to generate inputs that trigger as many distinct crashes as possible through the harness. A model's performance on each challenge is scored based on the number of distinct crash signatures it produces, capped at a predefined maximum and weighted by a difficulty coefficient. FuzzingBrain-Bench V1 consists of 77 challenges drawn from 43 open-source projects, with 36 C, 32 C++, and 9 Java/JVM challenges. We evaluate Claude Haiku 4.5, Claude Sonnet 4.6, and Claude Opus 4.8 on the full benchmark. Claude Opus 4.8 performs best, triggering crashes in 60 of 77 challenges and achieving a score of 196 out of 579. None of the three models triggers a crash in 13 challenges. The FuzzingBrain-Bench corpus and harnesses are publicly available at https://github.com/fuzzingbrain/FuzzingBrain-Bench.

发表机构

  • Texas A&M University(得克萨斯农工大学)
  • University of Novi Sad(诺威萨大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑