对自托管人工智能代理的自我状态攻击:操作系统防御能走多远?
Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?
浏览论文内容
中文总结 AI 辅助
研究自托管人工智能代理的自我状态攻击,刻画四轴攻击空间,通过收集活动跟踪数据实现攻击空间并评估防御策略,发现分层防御堆栈有效但有残余攻击面,为操作系统级防御研究提供新思路。
中文摘要 AI 辅助
自托管人工智能代理通过读写自身内存和配置文件来运行。代理可能因其自身状态受损而受到攻击,这种攻击可通过合法的操作系统系统调用实现,我们将这类威胁称为自我状态攻击。本文研究操作系统对这类攻击的恢复能力。我们正式刻画了一个四轴攻击空间(目标、机制、粒度、时间),研究预防、检测和恢复的结构限制,并引入基于工作负载的可检测性观点。为实例化该框架,我们收集了跨不同工作负载配置文件运行的代表性自托管代理的实时活动跟踪数据,将攻击空间实现为一个23单元矩阵、对真实自我状态文件的43个具体操作,并注入到这些跟踪数据中。然后我们评估了规范的和基于工作负载的防御策略。实证结果表明,分层防御堆栈(在指令和配置层进行访问控制预防、在内存层进行基于工作负载的检测以及定期备份用于恢复)在大多数攻击单元上有效,而在操作系统层面仍有一小部分残余攻击面在结构上难以区分。这些发现表明,针对新出现的自我状态攻击类别,需要重新考虑操作系统级防御,这可能为该领域开辟新的研究方向。
英文摘要
Self-hosted AI agents maintain persistent memory, instructions, and configuration that influence their future behavior. If an agent is compromised, an attacker can exploit the agent's legitimate write permissions to corrupt this self-state, making malicious and benign updates difficult to distinguish at the operating system (OS) level. We investigate how far existing OS mechanisms can prevent, detect, and recover from such self-state attacks. We formalize an attack space and evaluate representative OS defenses using four agent workloads and a Linux telemetry pipeline. Our results show a consistent limitation across defense dimensions. File-level controls either leave alternative mutation paths open or, when complete over the tested operations, also block corresponding legitimate updates. Detectors flag a substantial part of legitimate activity, while more selective methods cover only part of the attack space. Finally, protected backups successfully restore corrupted state, but require a trusted recovery point and may incur rollback cost. Overall, our results show that the main limitation is not OS observability. Indeed, the OS can enforce, observe, attribute, and recover self-state changes. Yet, generic OS defenses lack the decision context needed to combine broad operation coverage with selective decisions. Effective protection therefore requires self-state-aware mechanisms that exploit additional context beyond generic file and syscall behavior.
发表机构
- Center of Excellence for Generative AI(生成人工智能卓越中心)
- King Abdullah University of Science and Technology(卡瓦尔大学)
- The Swiss AI Lab, IDSIA-USI/SUPSI(瑞士人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。