发表机构
Meta(Meta)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现知识蒸馏对小型语言模型偏见存在非对称影响,明确任务下提升上下文遵循能力,模糊任务下破坏弃权校准,提出PCCD协议可捕捉聚合指标遗漏的危害。
AI 中文摘要
我们表明,小型指令调优语言模型中的知识蒸馏对偏见具有非对称影响。在明确任务(BBQ-disambig)上,来自Gemma-2-9B教师的基于响应的蒸馏提升了上下文遵循能力:对于最具偏见的基线模型(SmolLM2-1.7B-Instruct),它将上下文覆盖错误率从44%降至24%。在模糊任务(BBQ-ambig)上,相同的蒸馏破坏了逐项目弃权校准:15%的基线模型正确弃权的项目反而得到刻板印象答案,即便整体弃权率得以保留。该模式在第二个学生模型家族(OLMo-2-1B-Instruct)上重现,其中沉默损失为8%,填充沉默占新偏见的89%。在全部28种配置网格中,沉默损失与填充沉默的幅度无相关性(Spearman ρ=0.19,无统计学意义),表明两种效应源于不同机制。聚合刻板印象指标(CrowS-Pairs、整体BBQ刻板印象依赖评分)对两种效应取平均,掩盖了逐项目危害。我们将校准损失追溯至数据侧机制:对四个训练语料库的审计发现<0.5%的“弃权为答案形式”。带有弃权注入的监督微调(SFT)要么破坏解析,要么过度校正为“ trivial-refuser”模式(弃权率99.8%,歧义消除准确率0.2%),聚合指标会称该模式完全校准。我们提出Per-Condition Calibration Diagnosis(PCCD),这是一种评估弃权校准、上下文遵循和能力保留的三步协议。PCCD可捕捉到聚合评估遗漏的非对称危害和 trivial-refuser 失败模式。
英文摘要
We show that knowledge distillation (KD) in small instruction-tuned language models has asymmetric effects on bias, and that measuring them correctly requires accounting for where refusal mass moves and what the parser can legitimately score. On unambiguous tasks (BBQ-disambig), response-based distillation from a Mistral-7B teacher genuinely improves context-following for the most context-biased baseline (SmolLM2-1.7B-Instruct): among committed (non-abstaining) answers, the rate of overriding correct context with a stereotype falls from 44.5% to 37.2%, with accuracy rising from 0.55 to 0.61. On ambiguous tasks (BBQ-ambig), the same distillation degrades conditional refusal: 15% of the cases where the baseline correctly abstained instead receive stereotype answers (silence-loss), and the distilled refusal pattern only weakly preserves the baseline's (Spearman rho=0.44). The harm reproduces, aggravated, on a second student family (OLMo-2-1B-Instruct): silence-loss reaches 49% and filled-silence accounts for 95% of new bias. Two apparently stronger results are artifacts. An unconditioned override metric reports a 44% -> 23% improvement under a Gemma-2-9B teacher that shrinks to 44.5% -> 39.8% once conditioned on committed answers: the model abstains on 43% of items and its accuracy collapses from 0.55 to 0.35. An apparent cross-condition independence reverses to a positive correlation (rho=0.58, p<0.01) on the valid 19-configuration grid once parser-invalid logit-KD configurations are excluded and the parser is corrected. Aggregate metrics (CrowS-Pairs, overall BBQ Stereotype Reliance Score) average over both effects and conceal the per-item harm. We propose Per-Condition Calibration Diagnosis (PCCD), a three-step protocol evaluating refusal-pattern preservation, committed-answer context-following, and capability preservation. No configuration in our grid passes all three steps.
Comments18 pages, 5 figures. Preprint