反事实约束条件化在线蒸馏用于多约束指令遵循
Counterfactual Constraint-Conditioned On-Policy Distillation for Multi-Constraint Instruction Following
浏览论文内容
中文总结 AI 辅助
提出CC-OPD反事实约束条件化在线蒸馏方法,通过消融教师条件中的每个约束构建令牌级塑形奖励,提升多约束指令遵循性能,在多个基准上取得最优平均结果,1.5B学生超越7B教师。
中文摘要 AI 辅助
多约束指令遵循要求模型在多个同时激活的约束下对查询做出响应。即使是强大的指令调优模型也仍然会常规性地违反其中一些约束。现有方法要么通过外部验证器或学习评分器提供的序列级或令牌级强化学习奖励来增强监督,要么使用针对单一全上下文教师的在线蒸馏(OPD),但随着更多约束同时激活,该教师的概率质量会被稀释。我们提出CC-OPD(反事实约束条件化在线蒸馏),它反转了蒸馏中标准监督生成的方向。CC-OPD不是用超出学生所见的信息来丰富教师,而是依次消融教师条件中的每个约束,并从由此产生的每个令牌概率差异中构建每个约束的信号。由此产生的每个令牌留一法对数似然偏移被求和、裁剪,并作为令牌级塑形项添加到普通OPD奖励中。所有塑形项均从冻结的教师获得,在蒸馏过程中无需外部验证器,并且当聚合偏移为零时,奖励等于普通OPD。在两个Qwen模型对和七个基准测试中,CC-OPD在所有评估的学生训练方法中实现了最高的平均值。使用CC-OPD训练的1.5B学生模型在MulDimIF基准测试上超越了其自身的7B强化学习训练的教师模型。
英文摘要
Multi-constraint instruction following requires a model to respond to a query under many simultaneously active constraints. Even strong instruction-tuned models still routinely violate some of them. Existing approaches either augment supervision with sequence- or token-level RL rewards from external verifiers or learned graders, or use on-policy distillation (OPD) against a single full-context teacher whose probability mass becomes diluted as more constraints become simultaneously active. We propose CC-OPD (Counterfactual Constraint-Conditioned On-Policy Distillation), which inverts the standard supervision-generation direction in distillation. Rather than enriching the teacher with information beyond what the student sees, CC-OPD ablates each constraint from the teacher's conditioning in turn, and constructs the per-constraint signal from the resulting per-token probability differentials. The resulting per-token leave-one-out log-likelihood shifts are summed, clipped, and added to the vanilla OPD reward as a token-level shaping term. All shaping terms are obtained from the frozen teacher, without an external verifier during distillation, and the reward equals vanilla OPD wherever the aggregate shift is zero. Across two Qwen model pairs and seven benchmarks, CC-OPD achieves the highest average among all evaluated student-training methods. A 1.5B student trained with CC-OPD surpasses its own 7B RL-trained teacher on the MulDimIF benchmark.
发表机构
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。