arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LLMVul:来自真实生产仓库的LLM生成C/C++函数漏洞标注数据集

LLMVul: A Vulnerability-Labeled Dataset of LLM-Generated C/C++ Functions from Real Production Repositories

Mohammad Farhad, Shuvalaxmi Dass

arXiv 2609.10945首次发表:更新:

AI 中文总结

LLMVul从真实生产仓库中挖掘并标注了21,430个LLM生成的C/C++函数,识别出1,540个易受攻击函数,涵盖17个CWE类别,为漏洞检测和安全评估提供了大规模真实数据。

AI 中文摘要

大型语言模型(LLM)越来越多地被用于生成和辅助软件开发,然而现有的漏洞数据集主要关注人类编写的代码或受控提示环境。这限制了研究LLM生成代码在现实软件项目中出现的安全弱点的能力。我们提出了LLMVul,一个从真实生产仓库中挖掘的LLM生成C/C++函数的漏洞标注数据集。我们利用提交元数据和与AI相关的作者身份证据等来源信号,挖掘了从2022年11月13日至2026年9月3日为期4年的GitHub上AI辅助开发活动。经过过滤和去重后,LLMVul包含来自226个仓库的21,430个独特的C/C++函数,以及仓库、提交、函数、来源和AI工具元数据。我们使用互补的静态分析和基于模式的技术集成来建立漏洞标签,并为确认的易受攻击函数分配常见弱点枚举(CWE)类别。为了评估标签可靠性,我们还进行了独立的人工注释,并使用Cohen's kappa($k=0.79$)衡量评分者间一致性。LLMVul包含1,540个集成易受攻击的函数,涵盖17个独特的CWE类别,提供了比现有面向漏洞的LLM代码基准更多的真实世界LLM生成的易受攻击C/C++函数。通过保留代码级漏洞标签和生成/来源元数据,LLMVul能够支持关于漏洞检测、LLM生成代码的安全性评估以及AI辅助软件开发中漏洞模式分析的可重复研究。LLMVul数据集在此https URL公开可用。

英文摘要

Large language models (LLMs) are increasingly used to generate and assist with software development, yet existing vulnerability datasets largely focus on human-written code or controlled prompting environments. This limits the ability to study security weaknesses in LLM-generated code as it appears in real-world software projects. We present LLMVul, a vulnerability-labeled dataset of LLM-generated C/C++ functions mined from real production repositories. We mine AI-assisted development activity from GitHub over a 4 year period, from November 13, 2022 to September 3, 2026, using provenance signals such as commit metadata and AI-related authorship evidence. After filtering and deduplication, LLMVul contains 21,430 unique C/C++ functions from 226 repositories, together with repository, commit, function, provenance, and AI-tool metadata. We establish vulnerability labels using an ensemble of complementary static-analysis and pattern-based techniques and assign Common Weakness Enumeration (CWE) categories to confirmed vulnerable functions. To assess labeling reliability, we additionally conduct independent manual annotation and measure inter-rater agreement using Cohen's kappa ($k=0.79$). LLMVul contains 1,540 ensemble-vulnerable functions spanning 17 unique CWE categories, providing substantially more real-world LLM-generated vulnerable C/C++ functions than existing vulnerability-oriented LLM code benchmarks. By preserving both code-level vulnerability labels and generation/provenance metadata, LLMVul enables reproducible research on vulnerability detection, security evaluation of LLM-generated code, and analysis of vulnerability patterns in AI-assisted software development. The LLMVul dataset is publicly available at https://doi.org/10.5281/zenodo.22668216.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑