基于Atomic Red Team技术的符号化攻击链生成:谓词表示粒度的实证研究
Symbolic Attack Chain Generation from Atomic Red Team Techniques: An Empirical Study of Predicate Representation Granularity
浏览论文内容
中文总结 AI 辅助
本研究对比九类与五类谓词粒度的AALM,发现粒度对攻击链有效性影响小,81.3%结果一致,高粒度仅提升规划论证的内部结构分辨率。
中文摘要 AI 辅助
自动化攻击链生成对现代网络安全至关重要,但随着攻击者行为不断扩展,手动构建难以规模化。尽管采用PDDL的经典AI规划为自动化该过程提供了形式化方法,但其依赖于将技术准确转换为符号谓词。当前最先进的系统如AURORA采用九类攻击动作链接模型(AALM),但这一特定粒度的必要性尚未得到验证。本研究探究谓词表示粒度对规划有效性、成本及保真度的影响。研究采用的流程为:大型语言模型(LLM)执行转换,Fast Downward引擎执行确定性推理,对比完整九类AALM与基于Atomic Red Team(ART)执行证据经实证推导的简化五类方案。来自16种技术语料库的结果显示,规划有效性与成本对粒度基本不敏感,两类方案的结果相似度达81.3%。研究发现,更高粒度主要提升规划论证的内部结构分辨率,而非生成攻击链本身的可行性。
英文摘要
Automated attack chain generation is critical for modern cybersecurity, yet manual construction fails to scale as adversary behaviors expand. While classical AI planning using the Planning Domain Definition Language (PDDL) offers a formal method to automate this process, it relies on the accurate translation of techniques into symbolic predicates. Current state-of-the-art systems like AURORA employ a nine-category Attack Action Linking Model (AALM), but the necessity of this specific granularity remains unvalidated. This work investigates whether AURORA's nine-category taxonomy provides representational distinctions beyond those captured by a reduced, empirically derived scheme. Utilizing a pipeline where a Large Language Model (LLM) performs translation and the Fast Downward engine performs deterministic reasoning, the study compares the full nine-category AALM against a reduced five-category scheme derived empirically from Atomic Red Team (ART) execution evidence. Because the nine-category domain is constructed as a relabeling of the five-category domain, plan validity and cost are held identical between schemes by design; the substantive test of granularity's effect lies instead in the resulting predicate category resolution. There, a controlled A/B test isolates a case where a coarser scheme's plan passes every validity check while remaining operationally wrong: holding administrator privilege and being able to exercise it over a network logon prove to be causally distinct system states. Results from a sixteen-technique corpus show 81.3% identical plan outcomes across both schemes by construction, with a genuine predicate category resolution gain confined to a single technique out of sixteen. The findings suggest that higher granularity primarily enhances the internal structural resolution of a plan's justification rather than the viability of the generated attack chain itself.