CyberLLM:面向汽车网络安全的多智能体大语言模型框架,用于自主检测与受管控响应
CyberLLM: A Multi-Agent LLM Framework for Autonomous Detection and Guarded Response in Automotive Cybersecurity
浏览论文内容
中文总结 AI 辅助
CyberLLM是由大语言模型编排的多智能体框架,结合确定性检测层与大语言模型精化,在安全防护下实现汽车漏洞自主检测与修复,在基准测试中覆盖约70%漏洞且零误报,验证了LLM智能体自主防御的可行性。
中文摘要 AI 辅助
软件定义汽车(SDV)扩大了汽车在源代码、运行时日志和部署拓扑方面的攻击面,而安全约束禁止自主智能体在无监督情况下采取行动。本文提出CyberLLM,这是一个由大语言模型编排的多智能体框架,可在正式的运行时安全防护下自主检测漏洞并执行修复措施。检测环节结合了确定性层(正则表达式规则、抽象语法树分析器和拓扑图检查)与大语言模型精化步骤,从而在高召回率基础上补充高精度推理。决策智能体汇总检测结果,以以人为中心的资产分类法对其进行标记,并选择分层响应,通过带符号的跨会话记忆和重新规划反馈逐步提升置信度。所有操作在允许运行前都需通过四个上下文安全属性及独立操作对齐预言机的验证,被拒绝的操作会触发升级和重新规划。对称攻击流水线生成并重放漏洞利用程序,以便在相同场景下对防御方和攻击方进行测试。在一个独立的、带真实标注的基准测试集上,该测试集包含9个原始汽车电子控制单元(ECU)模块,涉及C、C++和Rust语言,植入了47个分层漏洞及干净的控制代码;始终开启的确定性层覆盖了34%的标注漏洞且精度完美,加入基于真实标注的大语言模型精化和完整性步骤后,覆盖范围大致翻倍至约70%(F1值为0.83),同时在干净控制代码上未产生任何误报。结果表明,当大语言模型智能体被包裹在确定性、可审计的安全防护范围内时,可执行有用的自主网络防御。
英文摘要
Software-Defined Vehicles (SDVs) expand the automotive attack surface across source code, runtime logs, and deployment topologies, while safety constraints forbid autonomous agents from acting without oversight. This paper presents CyberLLM, a multi-agent, LLM-orchestrated framework that autonomously detects vulnerabilities and executes remediations under a formal, runtime safety guard. Detection combines a deterministic layer (regex rules, AST analyzers, and topology graph checks) with an LLM refinement pass, so a high-recall floor is complemented by high-precision reasoning. A decision agent aggregates findings, tags them with a human-centric asset taxonomy, and selects a tiered response, ratcheting its confidence with signed cross-session memory and re-planning feedback. Every action is validated against four contextual security properties and an independent action-alignment oracle before it is allowed to run, and refused actions trigger escalation and re-planning. A symmetric attack pipeline generates and replays exploits so both sides can be exercised on the same scenarios. On an independent, ground-truthed benchmark of nine original automotive ECU modules in C, C++, and Rust seeding 47 layered vulnerabilities plus clean controls, the always-on deterministic layer covers 34\% of the labeled vulnerabilities at perfect precision, and adding the grounded LLM refinement and completeness passes roughly doubles coverage to about 70\% (F1 $0.83$) while producing zero false positives on the clean controls. The results indicate that LLM agents can perform useful autonomous cyber-defense when wrapped in a deterministic, auditable safety envelope.