K-Bench:面向智能体部署的LLM遗忘基准
K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments
浏览论文内容
中文总结 AI 辅助
K-Bench提出在智能体部署下评估LLM遗忘的新基准,检查ReAct智能体的六个通道,发现现有基准低估泄露,且权重中的秘密难以移除。
中文摘要 AI 辅助
诸如TOFU和MUSE等遗忘基准通过读取模型的最终答案来认证遗忘效果,其中拒绝回答的模型即被视为已遗忘。我们表明,一旦模型被部署为智能体,这种模型级认证便不再有效。我们引入了K-Bench,一个在智能体部署下对LLM遗忘进行评分的基准。K-Bench检查ReAct智能体暴露的所有六个通道,包括其思维链(CoT)、工具调用和工具观察,以及生成的摘要。如果秘密出现在其中任何一个通道中,则该查询被视为泄露。每个实验将秘密放置在智能体的三个来源(权重、提示或检索存储)中的恰好一个中。K-Score针对每个来源分别计算,并且仅在智能体保持可用时才计入遗忘。清除答案通道并不能使秘密不可恢复。在结构化检索中,秘密逐字保留在工具观察通道中,且总体泄露率不变。当秘密位于提示或检索存储中时,TOFU和MUSE报告无泄露,而部署的智能体在22%至86%的查询中仍会泄露它。当秘密位于权重中时,所评估的二十种已发表方法均未证明能移除它,并且只有一种输入损坏干预在评估的观察者下实现了选择性遗忘。排名靠前的方法因基础模型而异。一种拒绝调优方法在未经验证的知识移除的情况下抵抗了评估的提取。
英文摘要
Unlearning benchmarks such as TOFU and MUSE certify forgetting by reading the model's final answer, where a model that refuses to answer already counts as having forgotten. We show that this model-level certificate does not transfer once the model is deployed as an agent. We introduce K-Bench, a benchmark that scores LLM unlearning under agentic deployment. K-Bench inspects all six channels a ReAct agent exposes, including its chain-of-thought (CoT), tool calls and tool observations, and elicited summary. A query counts as leaked if the secret appears in any of them. Each experiment places the secret in exactly one of the agent's three sources (the weights, the prompt, or the retrieval store). The K-Score is computed separately for each source and credits forgetting only when the agent remains usable. Clearing the answer channel does not make the secret unrecoverable. On structured retrieval, the secret stays verbatim in the tool-observation channel and the aggregate leak rate is unchanged. When the secret lives in the prompt or the retrieval store, TOFU and MUSE report no leakage, while the deployed agent still leaks it on 22--86\% of queries. When the secret is in the weights, none of the twenty evaluated published methods demonstrably removes it, and only an input-corruption intervention reaches selective forgetting under the evaluated observer. The top-ranked method changes across base models. A refusal-tuning method resists the evaluated extraction without verified knowledge removal.
发表机构
- University of Technology Sydney(悉尼科技大学)
- CSIRO(澳大利亚联邦科学与工业研究组织)
机构由 AI 辅助整理,请以论文原文为准。