大语言模型辅助的自动驾驶车辆中攻击者可及软件弱点的动态威胁分析
LLM-Assisted Dynamic Threat Analysis for Attacker-Reachable Software Weaknesses in Autonomous Vehicles
浏览论文内容
中文总结 AI 辅助
该研究针对自动驾驶车辆软件弱点的动态威胁分析难题,探究LLM辅助Autoware的自动化测试工件生成,发现构建集成是核心障碍,推理模型编译测试用例的表现优于代码专用模型。
中文摘要 AI 辅助
自动驾驶车辆依赖大型安全关键软件栈,其中可被对抗性输入触及的弱点可能影响转向、制动或其他控制决策。静态分析可识别候选位点,但要动态确认其可利用性需可执行的测试工件,而手动构建此类工件十分困难。本研究探究大语言模型(LLM)能否针对开源自动驾驶栈Autoware实现该过程自动化。我们对185个软件包开展编译器精度级静态分析,识别出1375条决策规则、2274个验证检查及482条输入到安全输出的流,基于此构建弱点分类体系并采样出740个可及位点。采用两款本地开源权重LLM、无静态上下文的消融模型及朴素模板基线,生成3700组工件集,这些工件集经 sanitizers 下的真实构建编译、通过编译器闭环反馈修复,可执行时进行模糊测试。主要结果为构建集成失败分类体系,显示80%的首次编译失败源于依赖布线而非程序逻辑。推理模型首次尝试编译了64%的测试用例,而代码专用模型仅为6%。推理模型仅通过大量桩代码实现了完整对象可编译性,其测试用例中不足半数到达模糊测试环节,且观测到的37个崩溃均源自桩代码而非Autoware。在预算范围内未动态确认任何候选弱点。这些结果表明,构建集成而非候选生成或模糊测试,是可靠的LLM辅助全自动驾驶车辆软件栈动态分析的主要障碍。
英文摘要
Autonomous vehicles depend on large safety-critical software stacks, where weaknesses reachable from adversarial inputs may affect steering, braking, or other control decisions. Static analysis can identify candidate sites, but dynamically confirming exploitability requires executable test artifacts that are difficult to construct manually. We investigate whether large language models (LLMs) can automate this process for Autoware, an open-source autonomous-driving stack. We perform compiler-precise static analysis across 185 packages, identifying 1,375 decision rules, 2,274 validation checks, and 482 input-to-safety-output flows, from which we derive a weakness taxonomy and sample 740 reachable sites. Two local open-weight LLMs, a no-static-context ablation, and a naive-template baseline generate 3,700 artifact sets, which are compiled against the real build under sanitizers, repaired through compiler-in-the-loop feedback, and fuzzed when executable. The main result is a build-integration failure taxonomy showing that 80% of first-shot compilation failures arise from dependency wiring rather than program logic. The reasoning model compiled 64% of harnesses on the first attempt, compared with 6% for the code-specialized model. Repair achieved full object-compileability for the reasoning model only through extensive stubbing; fewer than half of its harnesses reached the fuzzer, and all 37 observed crashes originated in stubbed code rather than Autoware. No candidate weakness was dynamically confirmed within budget. These results show that build integration, not candidate generation or fuzzing, is the primary barrier to reliable LLM-assisted dynamic analysis of full autonomous-vehicle software stacks.
发表机构
- The University of Alabama(阿拉巴马大学)
- Department of Civil, Construction & Environmental Engineering, The University of Alabama(阿拉巴马大学土木、建筑与环境工程学院)
- Department of Computer Science, The University of Alabama(阿拉巴马大学计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。