arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从共识与分歧中学习:基于少数轨迹对比的无监督在线策略自蒸馏

Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast

Jiaxin Guo, Yanwei Yue, Xuanbo Fan, Chunyu Yang, Yan Zhang

arXiv 2608.08764首次发表:更新:

AI 中文总结

本研究提出无监督在线策略自蒸馏框架CoDA,通过共识与分歧的双分支设计,在无需外部监督的情况下提升语言模型推理能力,在竞赛级数学基准上表现优于自生成基线。

AI 中文摘要

在线策略自蒸馏通过在学生实际访问的状态上查询教师来提升语言模型的推理能力。现有方法通过让教师接触特权上下文形成强大的信息不对称,但它们从根本上依赖外部监督(如标准答案或验证器)来构建这一优势。我们提出CoDA(共识与分歧对齐),这是一种完全无监督的框架,可仅从模型自身未标注 rollout 的潜在不确定性结构中生成可靠的特权信息。CoDA 提取两个互补信号:在正分支中,答案级共识识别出稳定的推理模式,该模式以冻结的自教师为条件,为新的学生轨迹提供密集的分布指导;但由于一致并不保证正确,仅正样本蒸馏存在将相关错误放大为虚假共识的风险。为打破这一有害反馈循环,CoDA 纳入利用分歧的负分支:将少数轨迹视为不稳定的替代方案,通过以参考为锚的 KTO 式校准目标进行温和惩罚。这种非配对二元反馈提供了强大的正则化,无需假设共识是绝对真值。在竞赛级数学基准上的实证评估表明,CoDA 显著提升了推理能力,优于自生成基线,且能有效稳定训练以抵御错误共识。

英文摘要

On-policy self-distillation improves language-model reasoning by querying a teacher on states actually visited by the student. Recent methods create a powerful information asymmetry by exposing the teacher to privileged context, yet they fundamentally rely on external supervision---such as gold solutions or verifiers---to construct this advantage. We introduce CoDA (Consensus and Disagreement Alignment), a fully unsupervised framework that creates reliable privileged information entirely from the latent uncertainty structure of a model's own unlabeled rollouts. CoDA extracts two complementary signals. In the positive branch, answer-level consensus identifies a stable reasoning mode, which conditions a frozen self-teacher to provide dense distributional guidance on fresh student trajectories. However, because agreement does not guarantee correctness, positive-only distillation risks amplifying correlated errors into a false consensus. To break this harmful feedback loop, CoDA incorporates a negative branch that exploits disagreement: minority trajectories are treated as unstable alternatives and gently penalized via a reference-anchored, KTO-style calibration objective. This unpaired binary feedback provides robust regularization without requiring the strong assumption that the consensus is the absolute ground truth. Empirical evaluations on competition-level mathematical benchmarks demonstrate that CoDA significantly improves reasoning, outperforming self-generated baselines and effectively stabilizing training against erroneous consensus.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑