Fisher 信息校准的基于反馈的在线策略自蒸馏方法用于大语言模型
Fisher-Informed Recalibration for Feedback-Based On-Policy Self-Distillation of LLMs
浏览论文内容
中文总结 AI 辅助
针对基于反馈的在线策略自蒸馏中优化不稳定和性能崩溃问题,提出双分支框架 FIRE,通过 Fisher 迹校准监督信号,分离反馈方向与步长,实现稳定蒸馏并保持下游性能。
中文摘要 AI 辅助
基于反馈的在线策略自蒸馏已成为一种有前景的方法,使基础模型(更具体地说是大语言模型(LLMs))能够在外部反馈下从自身输出中学习,其中单一模型同时充当教师和学生。然而,此类方法可能表现出不稳定的优化,容易在训练过程中导致性能崩溃。为解决这一局限,我们提出了 FIRE(Fisher 信息校准,Fisher-Informed REcalibration),一个双分支框架,在微调期间对正确和错误的在线策略输出所应用的监督进行重新校准。对于正确的响应,FIRE 用重新加权的在线策略 SFT 替代自蒸馏;而对于错误的响应,FIRE 识别出对教师诱导更新产生不成比例影响的反馈组件,并相应地对反馈条件目标进行重新校准。两个分支都受到部分由 softmax Fisher 迹导出的词元级半径的影响。FIRE 将反馈应使模型移动的方向与模型在该方向上应移动的距离分离开来,同时保持行为良好的反馈监督不变。我们的实验表明,FIRE 提供了显著更稳定的自蒸馏,同时保持强大的下游性能,特别是在标准反馈条件蒸馏变得不稳定的设置中。
英文摘要
Feedback-based on-policy self-distillation has emerged as a promising approach for enabling foundation models, more specifically Large Language Models (LLMs), to learn from their own outputs under external feedback, with a single model serving as both teacher and student. However, such methods can exhibit unstable optimization, conducive to performance collapse during training. To address this limitation, we propose FIRE (Fisher-Informed REcalibration), a dual-branch framework that recalibrates the supervision applied to correct and incorrect on-policy outputs during fine-tuning. For correct responses, FIRE replaces self-distillation with re-weighted on-policy SFT, while for incorrect ones FIRE identifies feedback components that disproportionately influence the teacher-induced update and recalibrates the feedback-conditioned target accordingly. Both branches are influenced by a token-level radius derived in part from a softmax Fisher trace. FIRE separates which direction feedback should move the model from how far the model should move in that direction, while leaving well-behaved feedback supervision unchanged. Our experiments demonstrate that FIRE provides substantially more stable self-distillation while maintaining strong downstream performance, particularly in settings where standard feedback-conditioned distillation becomes unstable.
发表机构
- Purdue University(普渡大学)
- Yonsei University(延世大学)
- University at Buffalo-SUNY(纽约州立大学布法罗分校)
机构由 AI 辅助整理,请以论文原文为准。