AI 中文总结
本研究提出ForgetBench基准测试,通过两种评估范式与统一分析框架,揭示现有LLMs在长期知识保留与泛化间难以平衡的问题,为未来模型记忆机制优化提供方向。
AI 中文摘要
大型语言模型(LLMs)展现出强大的知识获取与推理能力,但在多次更新下保留先前获取知识的能力仍未得到充分理解。现有评估范式主要聚焦于单步推理或静态知识编辑,无法捕捉持续模型修改过程中知识保留与退化的时间动态。本研究提出ForgetBench——一个旨在系统表征LLMs在持续知识编辑下遗忘行为的基准测试。ForgetBench引入两种互补评估范式,即基于概念的问答和基于场景的问答,以分离孤立事实保留与结构化关系知识保留。基于顺序编辑框架,我们构建时间有序的知识流并在多个编辑阶段评估模型行为。为定量分析长期保留动态,我们进一步引入统一评估框架,对知识随时间的演化进行建模,可测量时间衰减、保留强度及跨实例稳定性。在不同模型与编辑方法上开展的大量实验表明,现有方法无法在长期保留与泛化质量之间取得平衡。我们的研究结果强调,未来LLMs需要更鲁棒的记忆机制,以随时间有效获取、更新并保留知识。代码将在论文接收后发布。
英文摘要
Large language models (LLMs) have demonstrated strong capabilities in knowledge acquisition and reasoning, yet their ability to retain previously acquired knowledge under repeated updates remains insufficiently understood. Existing evaluation paradigms primarily focus on single-step reasoning or static knowledge editing, which fail to capture the temporal dynamics of knowledge retention and degradation during continual model modification. In this work, we propose ForgetBench, a benchmark designed to systematically characterize forgetting behavior in LLMs under continual knowledge editing. ForgetBench introduces two complementary evaluation paradigms, namely concept-based QA and scenario-based QA, to disentangle isolated factual retention from structured relational knowledge preservation. Building upon a sequential editing framework, we construct temporally ordered knowledge streams and evaluate model behavior across multiple editing stages. To quantitatively analyze long-term retention dynamics, we further introduce a unified evaluation framework that models knowledge evolution over time, enabling the measurement of temporal decay, retention strength, and cross-instance stability. Extensive experiments across diverse models and editing methods demonstrate that existing approaches fail to strike a balance between long-term retention and generalization quality. Our findings highlight the need for more robust memory mechanisms that can effectively acquire, update, and preserve knowledge over time in future LLMs. Code will be released upon acceptance.
Comments9 pages, 4 figures