arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无需重启的遗忘:有状态大语言模型智能体的执行状态遗忘

Forgetting Without Restarting: Execution-State Unlearning for Stateful LLM Agents

Chao Yao, Yangbo Wei, Zhen Huang, Junhong Qian, Chenle Chen, Shaoqiang Lu, Chen Wu, Lei He

arXiv 2609.04875首次发表:更新:

发表机构

Arizona State University; Eastern Institute of Technology, Ningbo(亚利桑那州立大学; 宁波东方理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对有状态LLM智能体的遗忘缺陷,提出Provenance-Guided Selective Replay方法,实现高效精确的执行状态遗忘,且性能接近完全重置,重新计算令牌数最多减少9倍。

AI 中文摘要

长期运行的大语言模型(LLM)智能体是有状态的:除了对话记录外,它们还会积累压缩摘要、明文记忆、待执行工具计划,以及在每个服务API下的键值(KV)缓存。然而,如今的“遗忘”操作仅会删除一条明文记忆记录,就停止操作,而保留所有源自被撤销信息的制品。我们将执行状态遗忘形式化:在收到遗忘请求后,智能体的行为必须如同从未观测过目标信息一般。我们将运行时建模为确定性转换系统,证明了目标前的轨迹前缀可免费与该反事实世界共享,目标后的后缀因无令牌级归因而不可避免地被污染,且精确遗忘至少需要T−τ+1次重新计算的转换,其中τ是目标信息的注入步骤。基于来源的选择性重放(Provenance-Guided Selective Replay)作为跨提示、压缩记忆和缓存的跨层合约达到了这一界限:来源图定位注入点,检查点恢复简化为裁剪KV缓存,经净化的重放再生反事实后缀。通过在三个智能体套件、九个基准和三个模型系列上开展的 elicitation、随机及无字符串行为测试进行审计,发现记忆删除的泄漏情况未变,基于指令的遗忘在elicitation测试下完全失效(探测泄漏值Leak@probes = 1.00),来源编辑仍有80%的剧集受被撤销偏好影响,而选择性重放的表现与完全重置无差异,但重新计算的令牌数最多减少9倍。

英文摘要

Long-running LLM agents are stateful: beyond the transcript they accrete compressed summaries, plaintext memory, pending tool plans, and, under every serving API, a KV cache. Yet today's "forget" operations delete a plaintext memory record and stop, leaving every artifact derived from the revoked information intact. We formalize execution-state unlearning: after a forget request, the agent must behave as if it had never observed the target. Modeling the runtime as a deterministic transition system, we prove that the pre-target trajectory prefix is shared with this counterfactual world for free, that the post-target suffix is irreducibly tainted without token-level attribution, and that exact unlearning requires at least $T-τ+1$ recomputed transitions, where $τ$ is the target's injection step. Provenance-Guided Selective Replay attains this bound as a cross-layer contract spanning prompt, compressed memory, and cache: a provenance graph locates the injection point, checkpoint restoration reduces to cropping the KV cache, and sanitized replay regenerates the counterfactual suffix. Audited with elicitation, stochastic, and string-free behavioral tests across three agent suites, nine baselines, and three model families, memory deletion leaves leakage unchanged, instruction-based forgetting collapses under elicitation (Leak@probes = 1.00), and source redaction still acts on a revoked preference in 80% of episodes, while selective replay is indistinguishable from a full reset at up to 9x fewer recomputed tokens.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑