迈向对大语言模型微调中记忆知识无法泛化原因的机理理解
Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning
浏览论文内容
中文总结 AI 辅助
研究大语言模型微调中记忆知识无法泛化的问题,用自修补技术监测知识渗透动态,发现与知识电路未对准假设一致,还设计启发式策略验证,跨域实验证明发现具稳健性。
中文摘要 AI 辅助
微调大语言模型以注入新知识面临关键挑战:大语言模型能快速记忆新事实,但无法将其用于下游推理任务。我们将这种失败形式化为‘知道—使用差距’,其特征是记忆与泛化之间的准确率差距和时间滞后。为理解此现象,我们用未见知识微调大语言模型,并使用名为自修补的新颖干预技术监测知识在内部的空间渗透动态。自修补识别出重新定位表示能显著改善失败泛化情况的激活位置。这些结果与知识电路未对准假设一致:记忆表示可能存在于内部,但可能未被路由到计算有效的层。为证明这一诊断发现的实用性,我们设计了一种简单启发式策略,在泛化失败中恢复了58%至75%的神谕余量。跨域进行实验以验证这一发现的稳健性。
英文摘要
Fine-tuning LLMs to inject new knowledge faces a critical challenge: LLMs can quickly memorize new facts, yet fail to use them for downstream reasoning tasks. We formalize this failure as the Knowing-Using Gap, characterized by an accuracy gap and a temporal lag between memorization and generalization. To understand this phenomenon, we fine-tune LLMs with unseen knowledge and monitor the spatial permeation dynamics of the knowledge internally with self-patching, an adaptation of activation patching that scans all layer pairs at every fine-tuning check-point. Self-patching identifies activation locations where relocating representations substantially improves failed generalization cases. These results are consistent with a knowledge-circuit misalignment hypothesis: memorized representations can exist internally but may not be routed to computation-effective layers. Experiments are done cross-domain for the robustness of this finding. Building on this diagnosis, we propose layer-wise representation self-distillation (LRSD) that aligns knowledge representation from late storage layer to middle layer. LRSD keeps improving generalization after fine-tuning saturates and nearly doubles multi-hop chaining accuracy on Qwen, while leaving memorization intact.
发表机构
- HKUST(GZ)(香港科技大学(广州))
- HKUST(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。