发表机构
Accenture(埃森哲)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出多智能体LLM框架,从需求文档和架构图生成攻击树,实现设计时安全审查,经评估优于基线。
AI 中文摘要
设计层面的安全弱点可能源于需求、信任假设、缺失的控制措施以及实现开始前的数据流。现有的安全实践往往在代码编写之后才发现这些问题。我们提出了一种多智能体LLM框架,用于从产品需求文档和架构图中进行设计时安全分析。该框架将架构图解析为图表示,生成误用和失败案例,构建攻击树,检查治理与合规性差距,推荐缓解措施,分配企业安全领域标签,并生成候选的修订架构建议以供专家审查。该框架在推理时不从常见弱点枚举(CWE)数据库中检索。相反,它分析系统行为、信任边界、组件交互和数据流假设。误用案例作为中间表示,将发现与系统组件和攻击路径联系起来,而验证和细化循环则过滤无依据的发现,并提高基础性、可追溯性和可操作性。我们在一个微软参考标记的威胁建模示例、标记的合成PRD-架构对以及两个开放式系统(Berty和Gas Town)上评估了该框架。参考标记的案例支持威胁恢复和可操作性分析,而开放式案例评估有效性、噪声、可追溯性、可操作性、冗余性和攻击树质量。结果表明,与单次和消融基线相比,架构信息驱动的、误用引导的推理提高了审查质量。关键词:LLM多智能体系统、设计时安全、威胁建模、漏洞发现、架构图、安全分析、误用案例推导、攻击树、迭代推理、安全治理。
英文摘要
Design-level security weaknesses can arise from requirements, trust assumptions, missing controls, and data flows before implementation begins. Existing security practices often identify these issues after code is written. We present a multi-agent LLM framework for design-time security analysis from product requirement documents and architecture diagrams. The proposed framework parses architecture diagrams into graph representations, generates misuse and failure cases, constructs attack trees, checks governance and compliance gaps, recommends mitigations, assigns enterprise security-domain tags, and produces a candidate revised architecture recommendation for expert review. The framework does not retrieve from Common Weakness Enumeration (CWE) databases at inference time. Instead, it analyzes system behavior, trust boundaries, component interactions, and data-flow assumptions. Misuse cases act as intermediate representations that link findings to system components and attack paths, while a validation and refinement loop filters unsupported findings and improves grounding, traceability, and actionability. We evaluate the framework on a Microsoft reference-labeled threat-modeling example, labeled synthetic PRD--architecture pairs, and two open-ended systems: Berty and Gas Town. The reference-labeled case supports threat-recovery and actionability analysis, while the open-ended cases evaluate validity, noise, traceability, actionability, redundancy, and attack-tree quality. Results show that architecture-informed, misuse-driven reasoning improves review quality compared with single-shot and ablation baselines. Keywords: LLM Multi-Agent Systems, Design-Time Security, Threat Modeling, Vulnerability Discovery, Architecture Diagrams, Security Analysis, Misuse Case Derivation, Attack Trees, Iterative Reasoning, Security Governance.