arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33639cs.AI

基于LLM智能体的轨迹遗忘

Trajectory Unlearning on LLM-based Agents

  • Illinois Institute of Technology(伊利诺伊理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Yingdan Shi, Ren Wang

AI总结:

针对LLM智能体在长时程任务中的不良行为轨迹,提出轨迹级遗忘问题及GiRPO方法,通过注入遗忘轨迹并隔离归一化统计量实现稳定遗忘,在ALFWorld和WebShop上验证了其有效性与任务效用。

AI中文摘要:

现有的大语言模型(LLM)遗忘研究主要集中于移除特定知识,例如有害事实、私人数据或受版权保护的内容。然而,随着LLM日益被部署为自主智能体,一个根本性但被忽视的问题浮现出来:除了抑制智能体所知道的内容之外,智能体不应通过其行动轨迹重现不良行为。在本工作中,我们引入了轨迹级遗忘,这是一种新的问题设定,旨在移除长时程智能体任务中的特定行动轨迹,而非事实知识。我们识别出两个将轨迹遗忘与知识遗忘区分开来的根本性挑战:(1)我们的遗忘目标是智能体所“做”的,而非其“说”的;(2)轨迹是顺序依赖的行动序列,不能在不丢失步骤间结构的情况下分解为孤立的提示-响应对。为应对这些挑战,我们提出了组注入相对策略优化(GiRPO),该方法将遗忘轨迹注入策略rollout组并赋予惩罚性奖励,同时隔离归一化统计量,从而产生稳定且有界的遗忘信号,不会破坏正常任务轨迹的梯度更新。我们从两个应用场景——家务任务(ALFWorld)和在线购物(WebShop)——构建了轨迹遗忘基准,并设计了三个互补的指标来评估遗忘质量和模型效用。在ALFWorld和WebShop上的实验表明,GiRPO能有效遗忘目标轨迹,同时保持任务成功率,在遗忘质量和任务效用两方面均优于现有的知识遗忘基线。

英文摘要:

Existing large language model (LLM) unlearning has focused primarily on removing specific knowledge, such as harmful facts, private data, or copyrighted content. However, as LLMs are increasingly deployed as autonomous agents, a fundamental yet overlooked problem emerges: beyond suppressing what an agent knows, an agent should not reproduce undesired behaviors through its action trajectories. In this work, we introduce trajectory-level unlearning, a new problem formulation that targets the removal of specific action trajectories in long-horizon agentic tasks, rather than factual knowledge. We identify two fundamental challenges that distinguish trajectory unlearning from knowledge unlearning: (1) our unlearning target is what the agent \emph{does}, not what it \emph{says}; and (2) trajectories are sequentially dependent action sequences that cannot be decomposed into isolated prompt-response pairs without losing inter-step structure. To address these challenges, we propose Group-injected Relative Policy Optimization (GiRPO), which injects forget trajectories into the policy rollout group with penalized rewards and isolates the normalization statistics, yielding a stable and bounded unlearning signal that does not corrupt gradient updates for normal task trajectories. We construct trajectory unlearning benchmarks from two application scenarios, household tasks (ALFWorld) and online shopping (WebShop), and design three complementary metrics for evaluating forgetting quality and model utility. Experiments on ALFWorld and WebShop demonstrate that GiRPO effectively unlearns target trajectories while preserving task success rates, outperforming existing knowledge-unlearning baselines on both forgetting quality and task utility.

↑