arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35390cs.LG

混合策略蒸馏的归纳反馈

Inductive Feedback for Mixed-Policy Distillation

Amir Moeini, Huaijiang Zhu, Daniel Havir, Shangtong Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对口头反馈在策略内蒸馏中偏好迁移与指导未充分利用的问题,提出基于概率确认框架的归纳反馈方法,并引入共享轨迹估计器,在知识型与智能体基准上超越现有方案。

中文摘要 AI 辅助

口头反馈能够识别错误并给出修正建议,即使在缺乏可靠程序化验证器的情况下,也能为语言模型的后训练提供丰富的监督信号。这类反馈通常由能力较强的模型生成,可用于在策略内蒸馏中调节教师模型,即训练学生模型以匹配教师模型在学生生成的轨迹上的预测。然而,这种方法可能会迁移教师模型中未被反馈驱动的偏好,同时使反馈中的大量指导信息未被充分利用。我们发现这两个问题均源于标准的策略内蒸馏目标,具体而言是其最小化的散度以及用作目标的分布。我们提出的方法同时解决了这两个限制。首先,为了将反馈所传达的信息与教师模型固有的偏好分离开来,我们将口头反馈视为关于在给定前缀下某个特定词元是否应作为下一个词元的假设的证据。随后,我们采用一种概率确认框架,该框架基于教师模型在接收反馈前后的预测,唯一地确定词汇表上的一个排序。利用与该排序一致的确认分数,我们在学生模型的信任区域内构建一个目标分布。其次,为了从学生轨迹可能未使用的指导中学习,我们推导出一个简单的共享轨迹估计器,用于估计学生与目标分布在轨迹上的对称散度,通过重要性加权在两个方向上重用学生和反馈条件下的教师轨迹。实证评估表明,我们的方法在基于知识的基准和智能体基准上均优于常见的策略内蒸馏方案以及一种近期的对比变体。

英文摘要

Verbal feedback can identify errors and prescribe corrections, providing rich supervision for language-model post-training even when reliable programmatic verifiers are unavailable. Such feedback, often generated by a capable model, can be used to condition the teacher in on-policy distillation, which trains the student to match the teacher's predictions on student-generated rollouts. However, this approach can transfer teacher preferences that the feedback did not motivate, while leaving much of the feedback's guidance unused. We find that both problems come from the standard on-policy distillation objective, specifically the divergence it minimizes and the distribution it uses as its target. Our proposed method addresses both limitations. First, to isolate the information conveyed by the feedback from the teacher's inherent preferences, we treat verbal feedback as evidence for or against the hypothesis that a particular token comes next at a given prefix. We then adopt a probabilistic confirmation framework which uniquely determines an ordering over the vocabulary based on the teacher's predictions before and after it receives feedback. Using a confirmation score consistent with this ordering, we construct a target distribution within a trust region of the student. Second, to learn from guidance that student rollouts can leave unused, we derive a simple shared-rollout estimator of a symmetric divergence between the student and target distributions over rollouts, reusing student and feedback-conditioned teacher rollouts in both directions through importance weighting. Empirical evaluations show that our method outperforms the common on-policy distillation recipe and a recent contrastive variant on knowledge-based and agentic benchmarks.

发表机构

  • Alquist Robotics
  • University of Virginia(弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

↑