arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.08393cs.AIcs.CL

迈向对大语言模型微调中记忆知识无法泛化原因的机理理解

Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning

Lu Dai, Ziyang Rao, Yili Wang, Hanqing Wang, Hao Liu, Hui Xiong

首次发表
浏览论文内容

中文总结 AI 辅助

研究大语言模型微调中记忆知识无法泛化的问题,用自修补技术监测知识渗透动态,发现与知识电路未对准假设一致,还设计启发式策略验证,跨域实验证明发现具稳健性。

中文摘要 AI 辅助

微调大语言模型以注入新知识面临关键挑战:大语言模型能快速记忆新事实,但无法将其用于下游推理任务。我们将这种失败形式化为‘知道—使用差距’,其特征是记忆与泛化之间的准确率差距和时间滞后。为理解此现象,我们用未见知识微调大语言模型,并使用名为自修补的新颖干预技术监测知识在内部的空间渗透动态。自修补识别出重新定位表示能显著改善失败泛化情况的激活位置。这些结果与知识电路未对准假设一致:记忆表示可能存在于内部,但可能未被路由到计算有效的层。为证明这一诊断发现的实用性,我们设计了一种简单启发式策略,在泛化失败中恢复了58%至75%的神谕余量。跨域进行实验以验证这一发现的稳健性。

英文摘要

Fine-tuning LLMs to inject new knowledge faces a critical challenge: LLMs can quickly memorize new facts, yet fail to use them for downstream reasoning tasks. We formalize this failure as the Knowing-Using Gap, characterized by an accuracy gap and a temporal lag between memorization and generalization. To understand this phenomenon, we fine-tune LLMs with unseen knowledge and monitor the spatial permeation dynamics of the knowledge internally with self-patching, an adaptation of activation patching that scans all layer pairs at every fine-tuning check-point. Self-patching identifies activation locations where relocating representations substantially improves failed generalization cases. These results are consistent with a knowledge-circuit misalignment hypothesis: memorized representations can exist internally but may not be routed to computation-effective layers. Experiments are done cross-domain for the robustness of this finding. Building on this diagnosis, we propose layer-wise representation self-distillation (LRSD) that aligns knowledge representation from late storage layer to middle layer. LRSD keeps improving generalization after fine-tuning saturates and nearly doubles multi-hop chaining accuracy on Qwen, while leaving memorization intact.

发表机构

  • HKUST(GZ)(香港科技大学(广州))
  • HKUST(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

↑