发表机构
The University of Alabama(阿拉巴马大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出LLMSec-AV框架,利用含AV安全知识的LLM在Autoware上发现弱点,恢复率达76%,优于传统静态分析器,并引入18类漏洞分类法提升可解释性。
AI 中文摘要
自动驾驶汽车依赖数百万行安全关键软件,然而通用分析器无法理解哪些代码会影响车辆运动。本研究探讨了具备显式自动驾驶汽车(AV)安全知识的大型语言模型(LLM)能否在弱点检测上超越基于规则的工具。我们从漏洞记录、安全公告和AV安全文献中开发了一个包含18个弱点类别的AV漏洞分类法,并将其集成到基于LLM的自动驾驶汽车安全分析(LLMSec-AV)中。在Autoware上评估时,该框架将770个翻译单元分解为4,673个函数,并在四种提示条件下分析了161个函数,这些条件涉及分类法上下文、从374个先前披露中检索以及多步分析。结果与从上游修复中挖掘的46个弱点位置以及基于标志数量匹配的排列基线进行了比较。CodeQL、Semgrep、cppcheck和Clang静态分析器评估了相同的代码,并为CodeQL和Semgrep添加了AV特定规则。生成的模糊测试工具使用AFL++和消毒器进行了测试。LLM条件恢复了46个已知弱点位置中的最多76%,优于传统分析器。CodeQL、Semgrep和Clang静态分析器未匹配任何弱点,而cppcheck尽管有1,301条警报,仅匹配了一个。无辅助提示实现了相似的检测性能,表明分类法并未驱动召回率。然而,分类法上下文将分配给弱点类别的发现比例从接近零提高到超过80%,改善了可解释性和分诊。18个类别中有6个无法直接表示为静态分析规则。LLMSec-AV引入了一种AV特定、机器可读的漏洞分类法用于弱点发现,并表明LLM可以通过识别和组织真实AV软件中的安全相关发现来补充传统分析器。
英文摘要
Automated vehicles rely on millions of lines of safety-critical software, yet general-purpose analyzers do not understand which code can affect vehicle motion. This study asks whether large language models (LLMs) with explicit automated-vehicle (AV) security knowledge improve weakness detection beyond rule-based tools. We developed an AV vulnerability taxonomy with 18 weakness classes from vulnerability records, security advisories, and AV-security literature, and integrated it into LLM-based Security Analysis for Automated Vehicles (LLMSec-AV). Evaluated on Autoware, the framework decomposed 770 translation units into 4,673 functions and analyzed 161 functions under four prompting conditions involving taxonomy context, retrieval from 374 prior disclosures, and multi-step analysis. Findings were compared with 46 weakness locations mined from upstream fixes and a flag-volume-matched permutation baseline. CodeQL, Semgrep, cppcheck, and the Clang Static Analyzer evaluated the same code, with AV-specific rules added to CodeQL and Semgrep. Generated fuzzing harnesses were tested using AFL++ and sanitizers. LLM conditions recovered up to 76% of the 46 known weakness locations, outperforming conventional analyzers. CodeQL, Semgrep, and the Clang Static Analyzer matched none, while cppcheck matched one despite 1,301 alerts. Unaided prompting achieved similar detection performance, showing that the taxonomy did not drive recall. However, taxonomy context increased the share of findings assigned to a weakness class from near zero to over 80%, improving interpretability and triage. Six of the 18 classes could not be directly represented as static-analysis rules. LLMSec-AV introduces an AV-specific, machine-readable vulnerability taxonomy for weakness discovery and shows that LLMs can complement conventional analyzers by identifying and organizing safety-relevant findings in real AV software.
Comments21 pages, 3 figures, 4 tables