arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CADENCE:通过覆盖自适应在线策略蒸馏缩小推理差距

CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation

Satyam Kumar, Saurabh Jha

arXiv 2607.16955首次发表:更新:

AI 中文总结

研究针对在线策略知识蒸馏的问题,提出CADENCE统一框架,其DRIFT机制结合六个扩展组件,能解决冷启动崩溃等问题。在GSM8K和MATH-500上实验,将大教师模型蒸馏成小学生模型,大幅提升性能,且无需数据中心规模硬件。

AI 中文摘要

在线策略知识蒸馏将推理从大型教师模型转移到紧凑的学生模型,但现有方法存在三种复合失败模式:冷启动崩溃,即新学生模型对教师模型偏好的令牌赋予接近零的概率;状态无关的散度调度,即仅基于时间的正向/反向 KL 插值忽略学生模型的覆盖状态;二元奖励稀疏性,即通过通过/失败信号丢弃部分正确轨迹的信息。我们提出了 CADENCE,一个针对每种情况都有针对性修复的统一框架。其 DRIFT 机制在学生采样轨迹上调度正向 KL 和反向 KL 替代目标的逐令牌凸混合。六个组件对其进行了扩展:COVA,一种覆盖自适应β调度,加速正向到反向的过渡;FTB,一种分叉令牌增强,通过全局归一化熵参考将梯度集中在高熵位置;CCD,一种密集奖励,为不正确但接近的轨迹添加数值接近部分分数;LAP,简洁优先的正确展开强化;EMR,一种熵匹配校准正则化器;BSD,一个自引导自蒸馏阶段。在 GSM8K 和 MATH-500 上(校正后的 512 令牌协议,5 个种子,报告标准差),CADENCE 将一个 15 亿参数的教师模型蒸馏成一个 5 亿参数的学生模型,在 GSM8K 上达到 69.8±0.5% 的 pass@1(从预训练的 48.7% 提高;缩小了教师模型差距的 63.2%),使用 30 亿参数教师模型时达到 72.1±0.4%(缩小了 76.2%),比最强的匹配计算标签使用基线(DRIFT+二元奖励)高出 4.4±0.7 分。所有实验均在一台 Apple Mac Studio(M 系列,64GB 统一内存)上运行,表明有原则的蒸馏无需数据中心规模的硬件即可达到强大的推理质量。

英文摘要

On-policy knowledge distillation transfers reasoning from large teachers to compact students, but existing approaches suffer three compounding failure modes: (i) cold-start collapse, where a fresh student assigns near-zero mass to teacher-preferred tokens; (ii) state-agnostic divergence scheduling, where time-only forward/reverse-KL interpolation ignores the student's coverage state; and (iii) binary reward sparsity, where pass/fail signals discard information from partially correct traces. We present CADENCE, a unified framework with a targeted fix for each. Its DRIFT mechanism schedules a per-token convex mixture of forward-KL and reverse-KL surrogate objectives on student-sampled trajectories (per-token surrogates, not sequence-level KL gradient estimators). Six components extend it: (A) COVA, a coverage-adaptive $β$ schedule accelerating the forward-to-reverse transition; (B) FTB, a forking-token boost concentrating gradient at high-entropy positions via a globally-normalized entropy reference; (C) CCD, a dense reward adding numerical-proximity partial credit for incorrect-but-close traces; (D) LAP, brevity-preferential correct-rollout reinforcement; (E) EMR, an entropy-matching calibration regularizer; (F) BSD, a bootstrapped self-distillation phase. On GSM8K and MATH-500 (corrected 512-token protocol, 5 seeds, reported std), CADENCE distills a 0.5B student from a 1.5B teacher to 69.8 $\pm$ 0.5% GSM8K pass@1 (from 48.7% pretrained; 63.2% of the teacher gap closed) and to 72.1 $\pm$ 0.4% with a 3B teacher (76.2% closed), beating the strongest matched-compute label-using baseline (DRIFT+binary reward) by +4.4 $\pm$ 0.7 points. All experiments run on a single Apple Mac Studio (M-series, 64GB unified memory), showing principled distillation reaches strong reasoning quality without datacenter-scale hardware.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑