arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

结构化但脆弱:大型语言模型(LLMs)在网络安全决策中的局限性

Structured but Fragile: On the Limits of LLMs in Cybersecurity Decision-Making

Pasquale Malacaria, Yunxiao Zhang

arXiv 2608.20966首次发表:更新:

发表机构

Queen Mary University of London; University of Exeter(伦敦玛丽女王大学; 埃克塞特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究探讨LLMs在网络安全决策中的表现,发现其在明确攻击图结构时能产生接近优化基线的策略,但随复杂度增加而脆弱,对提示敏感,且无法稳健应用结构化推理,为AI辅助安全决策系统设计评估提供参考。

AI 中文摘要

大型语言模型(LLMs)正越来越多地应用于网络安全工作流程中,但目前仍不清楚它们是否能进行结构化的安全推理,还是仅仅依赖于表面线索和先验知识。我们在从现实威胁场景(包括勒索软件、供应链入侵、云滥用、Kubernetes攻击、POS恶意软件以及ICS/OT入侵)衍生的攻击图防御选择背景下研究这一问题。在预算约束下,LLMs必须选择安全控制措施以最小化攻击者的成功率。我们将它们的策略相互比较,并与作为结构化推理规范参考的博弈论优化基线进行对比。结果显示,LLMs表现出条件性能力:当提供明确的攻击图结构时,它们通常会产生接近优化基线的连贯策略;然而,其能力具有脆弱性——LLMs的行为随图复杂度增加而愈发脆弱,且对框架高度敏感,微小的提示变化会显著改变排名,仅将一个较差策略重新标记为“最优”就能大幅提升其评估结果。我们还观察到形式风险与LLM判断之间存在非单调关系:最接近最优的策略未必被LLM评估者排名最高。为进一步探究推理能力,我们要求LLMs为同一优化问题生成求解器,尽管生成的实现恢复了正确的高层表述,但与专用求解器相比扩展性较差。总体而言,我们的研究结果表明,LLMs可在受控表示下近似结构化网络安全推理,但无法稳健应用,这对AI辅助安全决策支持系统的设计与评估具有重要意义。

英文摘要

Large language models (LLMs) are increasingly used in cybersecurity workflows, yet it remains unclear whether they can perform structured security reasoning or merely rely on superficial cues and prior knowledge. We study this question in the context of defence selection over attack graphs derived from real-world threat scenarios, including ransomware, supply-chain compromise, cloud abuse, Kubernetes attacks, POS malware, and ICS/OT intrusion. Given a budget constraint, LLMs must select security controls to minimise attacker success. We compare their strategies against each other and against a game-theoretic optimization baseline used as a normative reference for structured reasoning. Our results show that LLMs exhibit conditional competence. When explicit attack-graph structure is provided, they often produce coherent strategies close to the optimization baseline. However, their capabilities are fragile. LLM behaviour becomes increasingly fragile with graph complexity and is highly sensitive to framing. Small prompt changes can substantially alter rankings, and merely relabeling a poor strategy as ``optimal'' dramatically improves its evaluation. We further observe a non-monotonic relationship between formal risk and LLM judgement: strategies closest to the optimum are not necessarily ranked highest by LLM evaluators. To further probe reasoning ability, we ask LLMs to generate solvers for the same optimization problem. While the generated implementations recover the correct high-level formulation, they scale poorly compared to a purpose-built solver. Overall, our findings show that LLMs can approximate structured cybersecurity reasoning under controlled representations, but do not apply it robustly. This has important implications for the design and evaluation of AI-assisted security decision-support systems.

Comments31 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑