发表机构
University of Louisiana at Lafayette(路易斯安那大学拉法叶分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VulValidate利用LLM智能体协调动态分析工具,基于可执行证据审计函数级漏洞标签,修正了BigVul等数据集中的大量错误标签,并提升了检测器性能。
AI 中文摘要
可靠的学习型漏洞检测需要高质量的标签,然而由漏洞修复提交构建的数据集可能仅因函数被安全补丁修改而将其标记为易受攻击。我们提出VulValidate框架,该框架利用LLM智能体协调动态分析工具,并根据运行时反馈构建漏洞触发实验。给定一个已标记的函数及其修复补丁,VulValidate重建易受攻击和修复后的版本,选择合适的工具和执行路径,细化触发输入,并比较运行时行为以评估函数级归因。我们审计了BigVul、PrimeVul和DiverseVul中最初标记为易受攻击的全部35,849个实例。我们确认了20,510个(57.2%),纠正了6,819个标签(19.0%),留下7,981个已攻击但未决(22.3%),以及539个(1.5%)无法成功测量。在冲突解决和字节级去重后,修正版本包含15,890个不同的已确认易受攻击函数体。在对581个抽样决策的盲审中,专家共识支持90.0%–92.0%的确认和92.6%–99.0%的标签纠正。在固定模型参数的情况下,修正后的评估在BigVul和DiverseVul上均降低了所有五个测试检测器的F1分数;使用修正标签重新训练,在各自数据集上四个检测器的F1分数得到提升。我们还发布了可复用的VulValidate技能、修正数据集和可复现的证据,以促进未来的漏洞检测研究。
英文摘要
Reliable learning-based vulnerability detection requires high-quality labels, yet datasets built from vulnerability-fixing commits may label functions as vulnerable simply because they were changed by a security patch. We present VulValidate, a framework that uses LLM agents to coordinate dynamic analysis tools and construct vulnerability-triggering experiments from runtime feedback. Given a labeled function and its fixing patch, VulValidate reconstructs vulnerable and fixed revisions, selects suitable tools and execution paths, refines triggering inputs, and compares runtime behavior to assess function-level attribution. We audit all 35,849 instances originally labeled vulnerable in BigVul, PrimeVul, and DiverseVul. We confirm 20,510 (57.2%), correct 6,819 labels (19.0%), leave 7,981 attacked but undecided (22.3%), and cannot successfully measure 539 (1.5%). After conflict resolution and byte-exact deduplication, the corrected release contains 15,890 distinct confirmed vulnerable function bodies. In a blinded review of 581 sampled decisions, expert consensus supports 90.0%--92.0% of confirmations and 92.6%--99.0% of label corrections. With model parameters fixed, corrected evaluation lowers F1 for all five tested detectors on both BigVul and DiverseVul; retraining with corrected labels improves F1 for four of five detectors on each dataset. We also release a reusable VulValidate skill, corrected datasets, and reproducible evidence for future vulnerability-detection research.
Comments21 pages