AI 中文总结
本研究首次系统研究同策略蒸馏中的成员推断,提出Leaky方法,通过采样轨迹并比较词元概率差,在15个目标上平均AUROC达0.875,显著优于基线。
AI 中文摘要
同策略蒸馏(On-policy distillation, OPD)训练学生模型,使其在学生生成的轨迹上匹配教师模型的下一词元分布。然而,为OPD训练提供给教师的特权信息可能包含敏感数据。学生是否会泄漏蒸馏期间提供给教师的记录中的私有信息,这一问题仍未被充分理解。据我们所知,我们首次对这一设置中的成员推断进行了系统性研究。我们发现,新鲜的学生轨迹暴露了稀疏的成员信号,而这些信号往往是固定的参考答案损失所遗漏的。这些信号与因训练其他记录而导致的概率变化混合在一起。我们引入了Leaky方法,该方法从目标模型中采样新鲜轨迹,并将其词元对数概率与在未包含候选记录的情况下训练的一组匹配参考模型中的最大值进行比较。Leaky对产生的差值应用Leaky ReLU,保留正差值并对负差值进行降权,作为对非成员中偶然正差值的一种近似校正。在涵盖数学、医学问答和代码生成的十五个目标上,Leaky在所有评估的基线方法中表现最佳,平均AUROC达到0.875,而在主要评估中每个目标上最强的基线平均AUROC为0.614。在相同的采样轨迹上,最强基线的平均AUROC为0.826。这些结果表明,通过OPD训练的学生模型能够暴露用于教师监督的记录成员身份,即使固定的参考答案损失几乎不提供成员身份的证据。
英文摘要
On-policy distillation (OPD) trains a student to match a teacher's next-token distributions on student-generated trajectories. However, privileged information supplied to the teacher for OPD training may contain sensitive data. Whether the student leaks private information about the records supplied to the teacher during distillation remains poorly understood. To the best of our knowledge, we present the first systematic study of membership inference in this setting. We find that fresh student trajectories expose sparse membership signals that fixed reference-answer losses often miss. These signals are mixed with probability changes caused by training on other records. We introduce Leaky, which samples fresh trajectories from the target model and compares its token log-probabilities with the maximum across matched reference models trained without the candidate records. It applies Leaky ReLU to the resulting gaps, preserving positive gaps and downweighting negative gaps as an approximate correction for incidental positive gaps in non-members. Across fifteen targets spanning mathematics, medical question answering, and code generation, Leaky outperforms all evaluated baselines and achieves mean AUROC 0.875, compared with 0.614 for the strongest baseline on each target in the main evaluation. On the same sampled trajectories, the strongest baseline achieves mean AUROC 0.826. These results show that students trained through OPD can expose the membership of records used for teacher supervision, even when fixed reference-answer losses provide little evidence of membership.