发表机构
The Chinese University of Hong Kong, Shenzhen; University of Notre Dame; Johns Hopkins University; Northwestern University; MIT(香港中文大学(深圳); 圣母大学; 约翰斯·霍普金斯大学; 西北大学; 麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大语言模型智能体的工具介导恢复问题,提出两阶段的Agentic Tool Unlearning框架,在RWKU和MUSE数据集上实现了更好的遗忘与保留效用平衡。
AI 中文摘要
大语言模型(LLMs)正越来越多地被部署为工具增强型智能体,其响应可依赖工具调用和外部观测,而非仅依赖模型参数。这造成了LLM遗忘的评估不匹配:以往的遗忘方法可能会抑制直接的参数回忆,但智能体仍可通过网络搜索、检索或数据库查找等工具恢复相同的遗忘目标。我们将这种失效模式称为工具介导的恢复,并研究智能体工具遗忘,其目标是在保留对保留知识的正常工具使用的同时,减少参数回忆和工具介导的恢复。为应对这一挑战,我们提出智能体工具遗忘(Agentic Tool Unlearning,ATU),这是一个两阶段框架。第一阶段应用参数知识遗忘以抑制直接回忆,第二阶段在模拟的工具增强型环境中执行轨迹级强化学习,以惩罚目标导向的工具行为和最终答案泄漏。在RWKU和MUSE数据集上针对不同LLM架构的实验表明,ATU在目标遗忘和保留效用之间实现了更好的平衡,使遗忘在工具增强型智能体部署下更具鲁棒性。
英文摘要
Large language models (LLMs) are increasingly deployed as tool-augmented agents, where responses can depend on tool calls and external observations rather than model parameters alone. This creates an evaluation mismatch for LLM unlearning: previous unlearning methods may suppress direct parametric recall, but an agent can still recover the same forget target through tools such as web search, retrieval, or database lookup. We identify this failure mode as tool-mediated recovery and study agentic tool unlearning, which aims to reduce both parametric recall and tool-mediated recovery while preserving normal tool use for retained knowledge. To address this challenge, we propose Agentic Tool Unlearning (ATU), a two-stage framework. The first stage applies parametric knowledge unlearning to suppress direct recall, while the second stage performs trajectory-level reinforcement learning in simulated tool-augmented environments to penalize target-seeking tool behavior and final-answer leakage. Experiments on RWKU and MUSE across different LLM architectures show that ATU achieves a better balance between target forgetting and retained utility, making unlearning more robust under tool-augmented agent deployment.