发表机构
The University of Alabama(阿拉巴马大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出人工验证的IoT固件漏洞检测基准IoTVulBench,通过多维度评估发现领域匹配训练数据与课程设计可提升跨语料库泛化性能,为IoT安全应用提供基准与可部署配置。
AI 中文摘要
IoT固件漏洞检测仍面临诸多挑战,原因在于异构的固件生态系统、资源受限的平台,以及现有基准的局限性。许多数据集是合成的或通用的,且缺乏经人工验证、污染筛查的标注,这限制了关于跨训练源、模型架构和课程策略的跨语料库泛化的证据。为解决这一缺口,本文提出了IoTVulBench,这是一个用于跨语料库固件漏洞检测的人工验证基准。IoTVulBench-Core由GitHub仓库构建,经三名专家审稿人验证,并在污染筛查的保留目标上对五种模型架构、两种调优方法和三种课程策略进行了评估,同时开展了集成、蒸馏和鲁棒性分析。在匹配的单源数据集中,在IoTVulBench上训练的模型达到了最高的马修斯相关系数(MCC),为0.58,而PrimeVul为0.44,D2A为0.39。分阶段课程学习将MCC提升至0.69,而多样性优化集成达到了0.73,相比最强的参考比较器(静态分析器,MCC为0.31)提升了0.42,相比PrimeVul提升了0.29。在0.5%的误报率下,该模型仅漏检21%的漏洞,而最强比较器的漏检率为71%。它在标识符重命名下保留了86%的性能,并展现出良好的校准性和基本可信的解释。这些发现表明,领域匹配的训练数据和课程设计,而非仅模型规模,是固件漏洞检测泛化的关键驱动因素。该结果为未来研究提供了基准,并为实际IoT安全应用提供了可部署的配置。
英文摘要
IoT firmware vulnerability detection is constrained by ecosystem heterogeneity, resource-limited platforms, and benchmark quality limitations. Existing datasets are often synthetic or general-purpose and lack human-verified, contamination-screened annotations, leaving cross-corpus generalization across training sources, architectures, and curriculum design underexplored. In this study, we have introduced IoTVulBench, a human-verified benchmark for cross-corpus firmware vulnerability detection. IoTVulBench was built from GitHub repositories, validated by three expert reviewers, and evaluated on a contamination-screened held-out target across five architectures, two tuning methods, and three curriculum strategies, with ensemble, distillation, and robustness analyses. Models trained on IoTVulBench reached the highest Matthews Correlation Coefficient (MCC) among undersampling-matched single-source datasets, at 0.58 versus 0.44 for PrimeVul and 0.39 for D2A. Staged curriculum learning raised MCC to 0.69, and a diversity-optimized ensemble reached 0.73. This gain represents a 0.42 MCC improvement over the strongest reference comparator, a static analyzer with an MCC of 0.31, and a 0.29 MCC improvement over the strongest single-source dataset, PrimeVul. At a 0.5% false-positive rate, the model missed only 21% of vulnerabilities versus 71% for the comparator. The model also retained 86% of its performance under identifier renaming, with strong calibration. These results indicate that domain-matched training data and curriculum design, rather than model scale alone, drive generalization in firmware vulnerability detection, and yield both a benchmark and deployment-ready configurations for IoT security.
Comments7 pages, 1 Figure, 6 Tables