LLM 智能体可以轻易篡改自己的痕迹
LLM Agents Can Easily Tamper With Their Own Traces
另 3 家 · 查看机构详情
- ELLIS Institute Tübingen(图宾根ELLIS研究所)
- Max Planck Institute for Intelligent Systems(马克斯·普朗克智能系统研究所)
- Tübingen AI Center(图宾根人工智能中心)
- Exponential Security Labs(指数安全实验室)
- Snyk(Snyk公司)
- University of Tübingen(图宾根大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究揭示本地 LLM 智能体可轻易篡改自身执行痕迹,多数框架缺乏防护,建议采用独立拦截机制确保痕迹完整性,以防隐藏恶意行为。
中文摘要 AI 辅助
异步监控、事件调查和合规审计主要依赖智能体痕迹来重建发生的事情。这些分析假设 LLM 智能体不能篡改自己的执行痕迹。我们表明,本地 LLM 智能体,如 Claude Code、Codex、Antigravity、Open Code 和 Grok Build,未能强制执行这一边界。所有测试的框架,除了 Muse Code,都允许智能体在被要求时删除其痕迹,而不会触发监控护栏。我们还验证了外部攻击者可以利用这一漏洞诱导痕迹删除。最后,我们表明,当智能体试图提高其奖励时,痕迹篡改行为会自然出现在前沿模型中。我们建议从业者确保痕迹记录通过独立于智能体控制的拦截机制进行,即使在主机完全受损的情况下也能保持痕迹完整性。总体而言,我们的发现识别了智能体基础设施中痕迹完整性的具体失败,这可用于隐藏如策划或破坏等不一致行为。
英文摘要
Asynchronous monitoring, incident investigations, and compliance audits primarily rely on agent traces to reconstruct what happened. These analyses assume that LLM agents cannot tamper with their own execution traces. We show that local LLM agents such as Claude Code, Codex, Antigravity, Open Code and Grok Build fail to enforce this boundary. All tested harnesses, except Muse Code, allowed agents to delete their traces when asked, without triggering monitor guardrails. We also validate that external attackers can exploit this gap to induce trace deletion. Finally, we show that trace tampering behavior emerges naturally in frontier models, when agents try to improve their rewards. We advise practitioners to ensure trace logging happens through an independent interception mechanism outside of the agent's control, preserving trace integrity even in cases of full host compromise. Overall, our findings identify a concrete failure of trace integrity in agent infrastructure which can be used to conceal misaligned behaviors like scheming or sabotage.