CROP:基于反事实的任务相关性用于选择性在线策略蒸馏
CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation
浏览论文内容
中文总结 AI 辅助
针对选择性在线策略蒸馏中任务相关性表征不足的问题,提出基于反事实的CROP方法,可识别更有用的监督位置,在两种师生设置中分别提升性能1.92和2.96个百分点。
中文摘要 AI 辅助
在线策略蒸馏(OPD)在从学生语言模型当前策略采样的轨迹上对其进行监督,但对监督价值不等的响应标记分配了相同的权重。选择性在线策略(Selective OPD)通过根据响应标记的估计训练值对其进行非均匀分配监督来解决这一限制。然而,大多数现有标准主要关注优化需求,例如不确定性或师生分歧,而任务相关性(即监督是否与当前输入的语义内容相关)作为一个互补维度,仍未得到直接表征。为解决这一差距,我们引入了在线策略蒸馏的反事实相关性(CROP),其通过释义校准的反事实敏感度裕度来实现任务相关性。对于每个源提示,CROP构建一个经过验证的原始-释义-反事实三元组,保持学生生成的序列固定,并通过响应位置对任务相关条件变化的敏感度(由其对保留语义的改写的敏感度校准)来衡量每个响应位置。匹配选择对照实验表明,CROP比随机选择或最低相关性选择识别出更有用的监督位置,而组件比较证实了反事实敏感度和释义校准的价值。在两个师生设置中,CROP比最强的非CROP选择器分别将综合性能提高了1.92和2.96个百分点。这些结果支持任务相关性作为选择性在线策略蒸馏的互补标准,并确立CROP为一种模型内部、特定对比的令牌级监督分配方法。
英文摘要
On-policy distillation (OPD) supervises a student language model on trajectories sampled from its current policy, but assigns equal credit to response tokens with unequal supervision value. Selective OPD addresses this limitation by allocating supervision non-uniformly across response tokens according to their estimated training value. Most existing criteria, however, focus primarily on optimization need, such as uncertainty or teacher-student disagreement, while task relevance, namely whether the supervision is tied to the semantic content of the current input, remains less directly characterized as a complementary dimension. To address this gap, we introduce Counterfactual Relevance for On-Policy Distillation (CROP), which operationalizes task relevance through a paraphrase-calibrated counterfactual sensitivity margin. For each source prompt, CROP constructs a validated original-paraphrase-counterfactual triplet, holds the student rollout fixed, and measures each response position by its sensitivity to a task-relevant condition change calibrated by its sensitivity to a meaning-preserving rewrite. Matched selection controls show that CROP identifies more useful supervision positions than random or lowest-relevance selection, while component comparisons confirm the value of both counterfactual sensitivity and paraphrase calibration. Across two teacher-student settings, CROP improves aggregate performance by 1.92 and 2.96 points over the strongest non-CROP selector. These results support task relevance as a complementary criterion for selective OPD and establish CROP as a model-internal, contrast-specific method for allocating token-level supervision.
发表机构
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。