连接智能体AI安全中的各个点:跨维度威胁分类、评估成熟度与开放挑战
Connecting the Dots in Agentic AI Security: A Cross-Dimensional Threat Taxonomy, Evaluation Maturity, and Open Challenges
查看机构详情
- Sungkyunkwan University(成均馆大学)
- Commonwealth Scientific and Industrial Research Organisation (CSIRO)(联邦科学与工业研究组织(CSIRO))
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对智能体AI安全,提出跨维度威胁分类T={S,B,P,A},基于66项研究分析实证覆盖与评估成熟度,指出研究差距并提炼13个开放问题,以推动系统化安全评估。
中文摘要 AI 辅助
智能体AI将LLM安全从生成内容扩展到持久状态、自主行动、工具使用以及与人类和其他智能体的交互。现有的威胁分类往往强调单个维度,模糊了入口点、受影响组件和安全后果之间的联系。已知的威胁格局也与实证研究所展示的覆盖范围有所不同。通过对2022年至2026年间发表的66项研究进行结构化审查,我们引入了T={S, B, P, A},一种跨维度表示,将受影响的功能或系统表面{S}、交互或信任边界{B}、被违反的安全属性{P}以及实证检验的架构{A}联系起来。我们分析了22项基于工件的红队研究和11个代表性安全基准,以刻画实证覆盖范围和评估成熟度。在所选研究中,证据集中于提示/推理、记忆和工具介导的攻击,主要发生在单智能体设置中。持久性、人机交互、复杂多智能体、系统性和长视野威胁获得的覆盖较少。这些发现描述了所选语料库,而非确定所有实证研究中的差距。异构指标、有限的适应性防御评估、架构不平衡以及不完整的执行状态捕获进一步限制了比较和可复现性。我们提出了13个开放研究问题,以指导更系统、架构感知且可复现的智能体AI安全评估。
英文摘要
Agentic AI extends LLM security beyond generated content to persistent state, autonomous actions, tool use, and interactions with humans and other agents. Existing threat classifications often emphasize individual dimensions, obscuring connections among entry points, affected components, and security consequences. The known threat landscape also differs from the coverage demonstrated by empirical research. Through a structured review of 66 studies published from 2022 to 2026, we introduce T={S, B, P, A}, a cross-dimensional representation linking affected functional or system surfaces {S}, interaction or trust boundaries {B}, violated security properties {P}, and empirically examined architectures {A}. We analyze 22 artifact-backed red-teaming studies and 11 representative security benchmarks to characterize empirical coverage and evaluation maturity. Within the selected studies, evidence concentrates on prompt/reasoning, memory, and tool-mediated attacks, predominantly in single-agent settings. Persistent, Human--Agent, complex multi-agent, systemic, and long-horizon threats receive less coverage. These findings describe the selected corpus rather than establish gaps across all empirical research. Heterogeneous metrics, limited adaptive defense evaluation, architectural imbalance, and incomplete execution-state capture further constrain comparison and reproducibility. We derive 13 open research questions to guide more systematic, architecture-aware, and reproducible security evaluation of agentic AI.