发表机构
Purdue University; University of Texas at Dallas(普渡大学; 德克萨斯大学达拉斯分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究围绕自主智能体安全领域,梳理了单智能体、多智能体及宏观层面的各类安全挑战,提出安全需成为智能体架构等的可验证属性,为可信自主智能体部署提供路线图。
AI 中文摘要
自主智能体越来越多地用于在受操作约束、组织策略、监管要求和技术标准管控的环境中执行关键任务,其安全性并非由单个行动的正确性决定,而是取决于整体行为是否符合其运行系统的规则和不变量。基于大语言模型(LLM)的智能体自主性不断增强,且越来越多地在组织边界间委派任务,其安全防护已从单一挑战演变为覆盖整个智能体栈的广泛互联领域。在单智能体层面,提示词、记忆、检索知识和工具接口的不可信输入构成攻击面;在多智能体场景中,委派与通信带来身份、信任、能力控制和决策透明度相关挑战,而底层模型路由和执行控制平面仍易受操纵及未经验证的模型来源影响。最根本的挑战或许是行为约束:单个可允许的行动序列可能集体违反系统级约束和安全不变量。在更宏观层面,供应链完整性、来源可追溯性、问责制和端到端可观测性仍是未解决的问题。这些方向有一个共同原则:安全必须成为管控智能体行为的架构、协议和运行时的可验证属性,而非可选的指导层。梳理这些挑战为可信自主智能体部署提供了路线图。
英文摘要
Autonomous agents are increasingly used to execute consequential tasks in environments governed by operational constraints, organizational policies, regulatory requirements, and technical standards. Their safety is therefore determined not by the correctness of individual actions, but by whether their overall behavior remains consistent with the rules and invariants of the systems in which they operate. As large language model (LLM)-based agents become more autonomous and increasingly delegate tasks across organizational boundaries, securing them evolves from a single challenge into a broad and interconnected landscape spanning the entire agentic stack. At the single-agent level, untrusted inputs through prompts, memory, retrieved knowledge, and tool interfaces create attack surfaces. In multi-agent settings, delegation and communication introduce challenges related to identity, trust, capability control, and decision transparency, while the underlying model routing and execution control plane remains vulnerable to manipulation and to unverified model provenance. Perhaps the most fundamental challenge is behavioral containment: sequences of individually permissible actions may collectively violate system-level constraints and safety invariants. At the broader level, supply-chain integrity, provenance, accountability, and end-to-end observability remain largely open problems. A common principle unifies these directions: security must become a verifiable property of the architectures, protocols, and runtimes that govern agent behavior, rather than an optional layer of guidance. Charting these challenges provides a roadmap toward trustworthy autonomous agent deployment.
Comments6 pages. Accepted to the ACM AI Leadership Summit 2026 (Visionary Track)