知而不言:基于回忆锚定蒸馏的大语言模型监督微调中事实访问失败的预防
Knowing but Not Saying: Preventing Factual Access Failures in LLM SFT via Recall-Anchored Distillation
浏览论文内容
中文总结 AI 辅助
本研究针对大语言模型SFT中出现的事实访问失败问题,提出回忆锚定蒸馏(RAD)方法,在保留目标领域适配性的同时恢复了部分分布外召回率,揭示了基础模型软分布的关键保留作用。
中文摘要 AI 辅助
监督微调(SFT)会导致模型在目标领域之外的事实行为出现退化,这种退化常被称为灾难性遗忘,但开放性事实失败并不一定意味着底层事实已被遗忘。本研究识别出一种更具体的现象:事实访问失败——在领域SFT后,模型在受限评估下仍能识别或排序正确答案,却无法在闭卷生成中输出该答案。通过基准级对比、同事实多项选择与生成探针、失败模式分析,我们表明SFT诱导的事实退化既包含真正的错误答案生成,也包含表达层面的失败,如冗长、格式不匹配、精确匹配假象。为解决该问题,我们提出回忆锚定蒸馏(RAD),一种基于基础模型的自蒸馏目标,通过在未标记的分布外(OOD)文本上使适配模型与原始基础模型的软延续分布对齐,从而保留分布外生成行为。RAD无需OOD标准答案、外部评判或标记事实数据。在MedMCQA上微调的三个骨干模型中,RAD在保留目标领域适配性的同时,恢复了部分丢失的OOD召回率;与相同OOD文本的重放相比,RAD表明关键的保留信号是基础模型的软分布,而非仅额外的文本暴露。
英文摘要
Supervised fine-tuning (SFT) can degrade factual behavior outside the target domain. This degradation is often described as catastrophic forgetting, yet open-ended factual failures do not necessarily imply that the underlying facts have been erased. In this work, we identify a more specific phenomenon, factual access failure: after domain SFT, models can still recognize or rank the correct answer under constrained evaluation, while failing to produce it in closed-book generation. Through benchmark-level comparisons, same-fact multiple-choice and generation probes, and failure-mode analysis, we show that SFT-induced factual degradation reflects both genuine wrong-answer generations and expression-level failures such as verbosity, formatting mismatch, and exact-match artifacts. To address this problem, we introduce Recall-Anchored Distillation (RAD), a base-anchored self-distillation objective that preserves out-of-distribution generation behavior by aligning the adapted model with the original base model's soft continuation distribution on unlabeled OOD text. RAD requires no gold OOD answers, external judges, or labeled factual data. Across three backbones fine-tuned on MedMCQA, RAD recovers a consistent portion of the lost OOD recall while preserving target-domain adaptation. Compared with replay on the same OOD text, RAD shows that the key preservation signal is the base model's soft distribution rather than additional text exposure alone.