arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

执行证明内存:通过验证实际发生的内容防御大语言模型智能体的伪造推理攻击

Proof-of-Execution Memory: Defending LLM Agents Against Forged-Reasoning Attacks by Verifying What Actually Happened

Md Habibur Rahman, Jaeho Kim

arXiv 2608.16032首次发表:更新:

AI 中文总结

该研究针对大语言模型智能体的伪造推理攻击,提出执行证明内存(PoEM)防御方案,通过维护防篡改的HMAC链式分类账验证实际执行步骤,使攻击成功率降至0%,且误报率极低、开销微小。

AI 中文摘要

大语言模型(LLM)智能体是无状态的,依赖外部内存来承载步骤间的上下文。由于智能体将该内存视为可信,能够写入内存的攻击者可操纵其行为。FARMA攻击无需恶意命令即可实现这一点:它在智能体的推理内存中插入伪造条目,声称所需的安全步骤已完成,从而使智能体跳过该步骤。与FARMA一同提出的防御方案SENTINEL会根据固定的可疑措辞列表对条目进行评分;其作者指出,知晓该列表的攻击者可改写伪造内容以规避防御,并将此作为未解决的问题。我们表明该差距比所述的更严重:仅要求语言模型改写伪造内容的自动攻击者首次尝试即可规避SENTINEL,在所有测试模型上将其保护效果降至零。我们还发现了一个能力悖论:该攻击在更强的模型上成功率高得多(在GPT-4o和GPT-4o-mini上为98%-100%,在Llama-3.1-8B上为44%),因为能力更强的智能体会更忠实地遵循改写后的声明,因此威胁随能力提升而增大。我们提出执行证明内存(Proof-of-Execution Memory, PoEM),该方案根本不检查内存。PoEM会维护一个独立的、防篡改的HMAC链式分类账,记录实际执行的安全步骤,仅受信任的动作层可写入,仅当分类账确认实际执行时才允许跳过步骤。攻击者可修改内存内容,但无法伪造从未运行的步骤的分类账条目,因此改写内容不再有效。在三个模型和三个场景中,PoEM将攻击成功率降至0%,同时保留合法操作(9个单元中有8个的误报率为0%,第9个为1.7%,处于采样噪声范围内),而SENTINEL错误阻止33%-50%的合法操作。PoEM还能抵御针对自身的攻击,仅增加微秒级的开销,且在真实的LangChain智能体中无需修改即可工作,PoEM准确保护其所管控的决策。

英文摘要

LLM agents are stateless and rely on external memory to carry context between steps. Because agents treat that memory as trustworthy, an adversary who can write to it can steer their behavior. The FARMA attack does this with no malicious command: it inserts fabricated entries into the agent's reasoning memory claiming a required safety step is already done, so the agent skips it. SENTINEL, the defense proposed with FARMA, scores entries against a fixed list of suspicious wordings; its authors note that an attacker who knows the list can reword the forgery and evade it, and leave this open. We show the gap is worse than stated. An automated attacker that simply asks a language model to reword the forgery evades SENTINEL on its first try, reducing its protection to zero on every model tested. We also find a capability paradox: the attack succeeds far more often on stronger models (98-100% on GPT-4o and GPT-4o-mini) than on Llama-3.1-8B (44%), because more capable agents follow reworded claims more faithfully, so the threat grows with capability. We propose Proof-of-Execution Memory (PoEM), which does not inspect memory at all. PoEM keeps a separate, tamper-evident, HMAC-chained ledger of the safety steps that actually executed, writable only by the trusted action layer, and allows a skip only if the ledger confirms real execution. An attacker can change what memory says but cannot forge a ledger entry for a step that never ran, so rewording no longer helps. Across three models and three scenarios, PoEM drives attack success to 0% while leaving legitimate operation intact (0% false positives in eight of nine cells, 1.7% in the ninth, within sampling noise), whereas SENTINEL wrongly blocks 33-50% of legitimate operations. PoEM also withstands attacks aimed at itself, adds microseconds of overhead, and works unchanged in a real LangChain agent. PoEM protects exactly the decisions it gates.

Comments8 pages, 6 figures, 5 tables. Code: https://github.com/bithabib/ai_security

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑