arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13731cs.AI

可信智能体人工智能:威胁态势、防御架构与开放挑战的综合网络安全与系统综述

Trustworthy Agentic AI: A Comprehensive Cybersecurity and Systems Survey on Threat Landscapes, Defense Architectures, and Open Challenges

Seyedakbar Mostafavi

首次发表
浏览论文内容

中文总结 AI 辅助

本综述针对智能体AI的安全挑战,提出系统安全参考框架,涵盖威胁分析、零信任防御架构及治理映射,综合206项研究,建立六维可信度分类法。

中文摘要 AI 辅助

从被动基础模型向自主、目标导向的智能体人工智能系统的转变,通过耦合递归认知推理循环、持久记忆架构、实时工具执行平面和多智能体协作拓扑,引入了前所未有的能力。然而,赋予概率神经核心跨文件系统、网络和云基础设施的执行权限,瓦解了经典的安全边界:自然语言同时充当输入数据、内部控制代码和通信协议,暴露了一个图灵完备的爆炸半径,其中不可信数据代表可执行指令。本综述为可信智能体人工智能提供了一个全面的系统安全参考框架,综合了206项基础研究和监管标准。我们将通用智能体架构形式化为一个有状态的五元组,并建立一个六维可信度分类法,涵盖安全性、安全、隐私、可解释性、公平性和问责制。我们系统地分析了执行循环内部和交互平面上的威胁面,构建了一个多层零信任纵深防御架构,集成了双大语言模型隔离、基于能力的访问控制、内核eBPF探针和沙箱运行时,审查了标准化评估基准,并将技术控制映射到国际人工智能治理框架。

英文摘要

The transition from passive foundation models to autonomous, goal-directed agentic AI systems has introduced unprecedented capabilities by coupling recursive cognitive reasoning loops, persistent memory architectures, live tool execution planes, and multi-agent collaboration topologies. However, granting probabilistic neural cores execution authority across filesystems, networks, and cloud infrastructure dissolves classical security perimeters: natural language simultaneously serves as input data, internal control code, and communication protocols, exposing a Turing-complete blast radius where untrusted data represents executable instructions. This survey delivers a comprehensive systems-security reference framework for trustworthy agentic AI, synthesizing 206 foundational studies and regulatory standards. We formalize the general agent architecture as a stateful 5-tuple and establish a 6-dimensional trustworthiness taxonomy covering security, safety, privacy, explainability, fairness, and accountability. We systematically analyze threat surfaces across intra-execution loops and interaction planes, formulate a multi-layered zero-trust defense-in-depth architecture integrating Dual-LLM isolation, Capability-Based Access Control, kernel eBPF probes, and sandboxed runtimes, review standardized evaluation benchmarks, and map technical controls to international AI governance frameworks.

发表机构

  • Yazd University(亚兹德大学)

机构由 AI 辅助整理,请以论文原文为准。

↑