arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30864cs.CLcs.AI

对抗性黑盒在线策略蒸馏中的持久负样本

Persistent Negatives for Adversarial Black-Box On-Policy Distillation

Haixu Ma, Saad Lahrichi, Weiwei Li, Kevin Han, Weiqiang Wu, Peggy Yang, Dongzhuo Li, Ruiyi Li, Serena Li, Gedi Zhou, Mingze Gao, Abhishek Kumar, Xiangjun Fan, Lizhu Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对黑盒在线策略蒸馏中判别器负样本随策略更新漂移的问题,提出持久负样本对抗性蒸馏,通过历史比较稳定奖励估计,在多个基准上提升性能。

中文摘要 AI 辅助

黑盒在线策略蒸馏(OPD)旨在当教师提供采样响应而非词元概率时,通过学生自身的生成来改进学生。对抗性蒸馏提供了一条途径:它学习一个判别器,用于区分提示匹配的教师和学生响应,并将其得分作为策略奖励。然而,在每一步从最新学生中采样判别器负样本,会将所学奖励与一个在每次策略更新后都会变化的负分布耦合在一起。我们通过持久负样本对抗性蒸馏来解决这一移动目标问题,这是一种实时池方法,用历史性的、提示匹配的教师-学生比较替换每个判别器批次中的一部分。在匹配的判别器计算量下,历史比较训练判别器,而GRPO则通过新鲜的学生响应保持在线策略。我们的分析将贝叶斯最优奖励识别为教师与负样本的对数密度比,并在明确假设下,展示了持久负样本如何锚定判别器并相对于新鲜负样本训练降低奖励估计的均方误差。在两个学生家族、三个评判者和四个评判聊天基准上,持久负样本对抗性蒸馏在匹配的判别器计算量下始终优于当前方法。它还产生了更平滑的新鲜策略判别器轨迹,减少了低于随机水平的下降。这些发现将判别器的负样本分布确定为黑盒在线策略蒸馏中的一个重要设计轴。

英文摘要

Black-box On-Policy Distillation (OPD) seeks to improve a student from its own generations when the teacher provides sampled responses but not token probabilities. Adversarial distillation offers one route: it learns a discriminator over prompt-matched teacher and student responses and uses its score as the policy reward. However, sampling discriminator negatives from the latest student at each step couples the learned reward to a negative distribution that changes after every policy update. We address this moving-target problem with persistent-negative adversarial distillation, a live-pool method that replaces a fraction of each discriminator batch with historical, prompt-matched teacher--student comparisons. Under matched discriminator compute, historical comparisons train the discriminator, while GRPO remains on-policy with fresh student responses. Our analysis identifies the Bayes-optimal reward as a teacher-to-negative log-density ratio and, under explicit assumptions, shows how persistent negatives anchor the discriminator and reduce reward-estimation MSE relative to fresh-negative training. Across two student families, three judges, and four judged-chat benchmarks, persistent-negative adversarial distillation consistently improves performance over current methods at matched discriminator compute. It also yields smoother fresh-policy discriminator trajectories, with fewer below-chance dips. These findings identify the discriminator's negative distribution as an important design axis in black-box on-policy distillation.

发表机构

  • Meta AI
  • University of Missouri(密苏里大学)

机构由 AI 辅助整理,请以论文原文为准。

↑