发表机构
University of Science and Technology of China; Zhejiang University(中国科学技术大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对在线对抗黑盒蒸馏中GRPO优势构建的目标不匹配问题,提出GRGC两阶段框架,通过高斯分组最优传输校准批评者奖励几何和策略侧组功率调制,提升优势构建的有效性,实验验证其有效性且开销可忽略。
AI 中文摘要
黑盒蒸馏是一种实用的途径,用于将仅暴露文本输出的API可访问大型语言模型的能力迁移到较小的学生模型中。最近的在线对抗方法(如GAD)通过在与学生之间形成对抗循环来改进SeqKD,其中批评者为基于GRPO的学生策略优化提供奖励,该优化针对学生采样的响应进行。然而,GRPO从同一提示的学生样本的组内相对奖励计算优势,而批评者主要被训练用于区分教师响应与学生响应。这种目标不匹配可能导致奖励组出现坍缩尺度或脆弱边际,从而产生脆弱的组优化信号。我们提出了分组奖励几何条件化(GRGC),这是一个两阶段框架,通过在批评者训练和策略优化期间塑造学生侧奖励组来改进优势构建。为了改善批评者侧条件化,高斯分组最优传输校准在训练期间正则化批评者,通过将排序的提示级奖励匹配到以组为中心的高斯分位数,产生具有非坍缩扩展和平滑秩间差距的奖励组。基于这种条件化的奖励几何,策略侧组功率调制在奖励组转换为优势之前重塑提示级奖励组,保留批评者诱导的排序,同时增加优化相关的边际可分性。跨不同教师、学生模型家族和规模以及训练数据集的广泛实验证明了GRGC在分布内和分布外评估中的有效性,同时相对于GAD引入可忽略的开销。代码可在该https URL获取。
英文摘要
Black-box distillation is a practical route for transferring capabilities from API-accessible large language models that expose only text outputs into smaller student models. Recent on-policy adversarial methods such as GAD improve over SeqKD by forming an adversarial loop between a critic and a student, where the critic provides rewards for GRPO-based student policy optimization over the student's sampled responses. However, GRPO computes advantages from the within-group relative rewards of student samples for the same prompt, whereas the critic is trained primarily to distinguish teacher responses from student responses. This objective mismatch can produce reward groups with collapsed scale or fragile margins, leading to brittle grouped optimization signals. We propose Groupwise Reward Geometry Conditioning (GRGC), a two-stage framework that improves advantage construction by shaping student-side reward groups during both critic training and policy optimization. To improve critic-side conditioning, Gaussian groupwise Optimal Transport calibration regularizes the critic during training to produce reward groups with non-collapsed spread and smooth rank-wise gaps by matching sorted prompt-wise rewards to group-centered Gaussian quantiles. Building on this conditioned reward geometry, policy-side group power modulation reshapes the prompt-wise reward groups before they are converted into advantages, preserving the critic-induced ordering while increasing optimization-relevant margin separability. Extensive experiments across diverse teachers, student model families and scales, and training datasets demonstrate the effectiveness of GRGC on both in-distribution and out-of-distribution evaluations, while introducing negligible overhead over GAD. The code is available at https://github.com/2018cx/GRGC.
CommentsNeurIPS 2026