用于可组合非结构化知识编辑的混合策略自编辑
Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing
- University of Tennessee(田纳西大学)
- Rutgers University(罗格斯大学)
- Purdue University(普渡大学)
- University at Albany(奥尔巴尼大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对现有非结构化知识编辑存在的可组合性缺失问题,提出HPSE方法,通过构建混合生成结果实现主动自蒸馏,在多种LLM及编辑工具上取得即插即用的性能提升。
AI中文摘要:
大型语言模型(LLMs)在各类自然语言任务中表现出色,但它们基于静态语料库进行训练,在快速变化的世界中其知识会迅速过时。这推动了知识编辑(KE)的发展,该技术旨在更新LLM中的特定知识,同时不改变其他无关知识。近期研究已从结构化知识三元组转向非结构化知识编辑(UKE),其中编辑内容是一段自由形式的文本,可能同时陈述多个事实。然而,现有的编辑工具虽能注入此类文本,却未能有效利用它:编辑后的模型能够回忆起该文本,既无法回答关于其事实的原子级问题,也无法将这些事实组合成多跳推理。我们将这一缺失的属性(称为可组合性)归因于编辑工具被动依赖固定文本作为唯一学习源。为此,我们将知识编辑建模为从同一模型的特权上下文状态进行主动自蒸馏,无需外部监督。我们进一步发现,由于注入知识的新颖性,编辑前模型的自身生成结果很少覆盖该知识,这限制了纯在线策略蒸馏的有效性。为弥补这一差距,我们提出HPSE(混合策略自编辑),该方法构建混合生成结果,在学生模型自身覆盖失败的位置,将缺失的事实精准放置在其轨迹上,而在其他位置保持在线策略。我们从理论上分析了HPSE相比纯在线策略蒸馏的优势,并在四种LLM主干模型和两种KE编辑工具上,在各类场景下通过实验证实了其即插即用的改进效果。
英文摘要:
Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world. This motivates knowledge editing (KE), which updates specific knowledge in an LLM without changing unrelated others. Recent works move from structured knowledge triples toward unstructured KE (UKE), where the edit is a free-form passage that may state multiple facts at once. Nonetheless, existing editors inject such a passage yet fail to use it: the edited model can recall the passage, but can neither answer atomic questions about its facts nor compose them into multi-hop reasoning. We attribute this missing property, which we term composability, to editors' passive reliance on the fixed passage as the sole learning source. In response, we cast editing as a proactive self-distillation from a privileged in-context state of the same model, which requires no external supervision. We further reveal that due to the novelty of the injected knowledge, the pre-edited model's own rollouts rarely cover it, which limits the effectiveness of pure on-policy distillation. To close this gap, we propose HPSE, which builds a hybrid rollout that steps in to place missing facts onto the student's own trajectory precisely where its coverage fails, while staying on-policy elsewhere. We theoretically analyze HPSE's advantage over pure on-policy distillation, and empirically establish its plug-and-play improvements across four LLM backbones and two KE editors under various scenarios.