发表机构
Indian Institute of Information Technology, Kalyani; RYVANE(卡利亚尼印度信息技术学院; RYVANE)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过大规模红队测试证明智能体记忆可被用作跨会话隐蔽隐写信道,发现25.4%的试验可恢复载荷,并主张将记忆完整性、信息流控制和隐蔽信道检测纳入智能体系统安全要求。
AI 中文摘要
智能体记忆能否抵御隐蔽的隐写攻击?我们开展了一项大规模红队演练,以测试智能体是否能在一次会话中编码攻击者控制的字符串,并在另一次会话中恢复这些字符串,同时不触发安全监督。遵循SHADE-Arena风格的任务,我们在原本良性的任务中嵌入恶意侧任务,利用隐写术编码秘密字符串,并使用独立的任务完成度与安全监督进行评估。我们测试了14,000次攻击试验,涵盖91种模型-攻击配置、13个模型和7种隐写方案:藏头诗(Acrostic)、藏中诗(Mesostic)、藏尾诗(Telestich)、词汇选择(Word-choice)、大小写(Capitalization)、位置(Position-based)和同形字(Homoglyph)。41.2%的试验成功在生成的响应中编码目标载荷,29.8%的试验在记忆写入步骤后仍保持编码完整。总体而言,25.4%的试验在检索时产生可恢复的载荷,20.1%实现精确匹配恢复,包括合成API密钥和从未暴露给接收方的加密材料。在精确匹配传输中,71.4%同时通过主任务完成度评分,并被监督模型独立判定为安全,这对应于所有试验的14.4%,即在这些试验中,成功的隐蔽传输在任务级评估下看似普通、良性的交互。我们的结果表明,智能体记忆可以充当持久的跨会话隐蔽信道。结果进一步表明,主要瓶颈出现在记忆持久性而非检索:一旦隐写载荷通过记忆写入阶段,相当一部分仍可恢复。因此,我们认为记忆完整性、信息流控制和隐蔽信道检测应成为智能体系统的明确安全要求。
英文摘要
Is agentic memory robust against stealthy steganographic attacks? We carry out a large-scale red-teaming exercise to test whether agents can encode attacker-controlled strings in one session and recover them in another without triggering safety oversight. Following SHADE-Arena-style tasks, we embed malicious side tasks to encode secret strings using steganography within otherwise benign tasks and evaluate them using independent task-completion and safety oversight. We test 14,000 attack trials spanning 91 model-attack configurations, 13 models, and 7 steganographic schemes: Acrostic, Mesostic, Telestich, Word-choice, Capitalization, Position-based, and Homoglyph. 41.2% of trials successfully encode the target payload in the generated response, while 29.8% survive the memory-write step with the encoding intact. Overall, 25.4% of trials yield a recoverable payload at retrieval, and 20.1% achieve exact-match recovery, including synthetic API keys and cryptographic material never exposed to the receiver. Among the exact-match transmissions, 71.4% also pass primary task-completion scoring and are independently judged safe by the oversight model, corresponding to 14.4% of all trials in which a successful covert transmission would appear to be an ordinary, benign interaction under task-level evaluation. Our results demonstrate that agentic memory can function as a persistent cross-session covert channel. The results further show that the principal bottleneck occurs at memory persistence rather than retrieval: once a steganographic payload survives the memory-write stage, a substantial fraction remains recoverable. We therefore argue that memory integrity, information-flow control, and covert-channel detection should be explicit security requirements for agentic systems.