arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

耦合校准与学习:在无目标域奖励反馈的LLM蒸馏中缓解教师偏差

Coupled Calibration and Learning: Mitigating Teacher Bias in LLM Distillation without Target-Domain Reward Feedback

Haichen Hu, Yuheng Zhang, David Simchi-Levi

arXiv 2609.17474首次发表:更新:

发表机构

MIT; UIUC(麻省理工学院; 伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出耦合校准与学习(CCL)算法,通过令牌级分支耦合教师校准与学生更新,仅用源反馈,在理论上证明可消除教师偏差并收敛至最优学生,无需目标域奖励。

AI 中文摘要

大语言模型(LLM)蒸馏旨在将强大教师模型的能力迁移至较小的学生模型。然而,直接模仿也可能迁移教师系统的偏差和错误。在协变量偏移下,当教师对目标问题的可靠性不确定且无法获得目标域奖励反馈时,这一挑战尤为突出。我们提出了耦合校准与学习(CCL),一种LLM蒸馏算法,通过令牌级分支将教师校准与学生更新耦合,仅使用源问题的奖励反馈。每次迭代利用源反馈校准教师,然后使用校准后的教师在目标问题上训练学生。更新后的学生反过来为后续校准提供信息。在自回归策略框架中,我们证明了输出学生与预言机学生的期望平均Kullback-Leibler散度以迭代次数的多项式速率收敛到零。预言机在学生类别内最大化真实的参考正则化目标奖励,该类别无需表示无约束的最优策略。我们的分析量化了投影学生梯度更新的进展,同时控制了教师校准中的误差。我们进一步建立了与正则化直接匹配的分离:即使教师实现了比每个学生策略更高的正则化目标奖励,其相对于预言机学生的误差仍可能保持远离零。这些结果表明,LLM蒸馏可以通过耦合校准与学习克服持续的教师偏差并恢复最优学生,而无需目标域奖励反馈。

英文摘要

Large language model (LLM) distillation aims to transfer the capabilities of a powerful teacher to a smaller student. Direct imitation, however, can also transfer the teacher's systematic bias and errors. This challenge is particularly pronounced under covariate shift, when the teacher's reliability on target questions is uncertain and target-domain reward feedback is unavailable. We propose Coupled Calibration and Learning (CCL), an LLM distillation algorithm that couples teacher calibration with student updates through token-level branching, using reward feedback only on source questions. Each iteration calibrates the teacher using source feedback and then uses the calibrated teacher to train the student on target questions. The updated student, in turn, informs subsequent calibration. In an autoregressive policy framework, we prove that the output student's expected average Kullback-Leibler divergence to the oracle student converges to zero at a polynomial rate in the number of iterations. The oracle maximizes the true reference-regularized target reward within the student class, which need not represent the unrestricted optimal policy. Our analysis quantifies the progress of projected student gradient updates while controlling the error in teacher calibration. We further establish a separation from regularized direct matching: its error relative to the oracle student can remain bounded away from zero even when the teacher achieves higher regularized target reward than every student policy. These results demonstrate that LLM distillation can overcome persistent teacher bias and recover the optimal student through coupled calibration and learning, without target-domain reward feedback.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑