arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

让数据决定:通过离策略蒸馏进行持续预训练中的监督分析、能力权衡和自适应目标路由

Let the Data Decide: Supervision Analysis, Capability Trade-offs, and Adaptive Objective Routing in Continued Pre-Training via Off-Policy Distillation

Jiangan Yuan, Zhixuan Li, Han Xu

arXiv 2607.16246首次发表:更新:

发表机构

Baidu Inc.(百度公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究离策略蒸馏中训练数据、目标参数化和模型能力的相互作用,通过分解问题进行“目标到能力”和“数据到目标”分析,引入诊断指标量化张力,探讨自适应目标路由,表明有效路由取决于信号质量,将持续预训练重构为监督设计问题。

AI 中文摘要

离策略蒸馏是大语言模型预训练的核心,但训练数据、目标参数化和模型能力之间的相互作用仍未得到充分表征。我们通过将此问题分解为两个问题来研究top-$k$-截断、温度缩放的离策略蒸馏:一是训练目标如何塑造token级监督和下游性能的“目标到能力”分析,二是数据异质性应如何为目标路由提供信息的“数据到目标”分析。我们首先表明语言建模目标($L_{\mathrm{LM}}$)和知识蒸馏目标($L_{\mathrm{KD}}$)会诱导出系统不同的能力配置文件,并将这种差异追溯到“直接观察token强化”和“教师支持的替代监督”之间的梯度级张力。为了量化这种张力,我们引入了诊断指标——支持覆盖率、观察token概率质量和教师分布集中度,并通过控制扫描表明支持大小$k$控制着覆盖率-清晰度权衡,而蒸馏温度控制支持内概率分配。然后我们研究了自适应目标路由:一种在数学和代码上应用$L_{\mathrm{LM}}$,在通用领域数据上应用$L_{\mathrm{KD}}$的领域级策略比两个单目标基线都有一致的收益,而基于观察token概率质量或教师熵的token级路由未能始终如一地匹配单目标基线。这些结果表明,有效的目标路由更多地取决于路由信号的质量而不是路由粒度,将通过离策略蒸馏的持续预训练重新构建为一个结构化的、数据条件监督设计问题,而不是一个全局超参数选择。

英文摘要

Off-policy distillation is now central to large language model pre-training, yet how training data, objective parameterization, and model capabilities interact remains poorly characterized. We studies top-$k$-truncated, temperature-scaled off-policy distillation by decomposing this problem into two questions: an \emph{objective-to-capability} analysis of how the training objective shapes token-level supervision and downstream performance, and a \emph{data-to-objective} analysis of how data heterogeneity should inform objective routing. We first show that the language-modeling objective ($L_{\mathrm{LM}}$) and the knowledge-distillation objective ($L_{\mathrm{KD}}$) induce systematically different capability profiles, and trace this divergence to a gradient-level tension between \emph{direct observed-token reinforcement} and \emph{teacher-supported alternative supervision}. To quantify this tension, we introduce diagnostic metrics -- support coverage, observed-token probability mass, and teacher-distribution concentration -- and show via controlled sweeps that the support size $k$ governs a coverage-sharpness trade-off, while distillation temperature controls within-support probability allocation. We then examine adaptive objective routing: a domain-level policy that applies $L_{\mathrm{LM}}$ to math and code and $L_{\mathrm{KD}}$ to general-domain data yields consistent gains over both single-objective baselines, whereas token-level routing based on observed-token probability mass or teacher entropy fails to consistently match the single-objective baseline. These results suggest that effective objective routing depends less on routing granularity than on the quality of the routing signal, reframing continued pre-training via off-policy distillation as a structured, data-conditional supervision-design problem rather than a global hyperparameter choice.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑