arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

作为无标签自蒸馏特权上下文的共识

Consensus as Privileged Context for Label-Free Self-Distillation

John Gkountouras, Josip Jukić, Ivan Titov

arXiv 2607.13643首次发表:更新:

发表机构

ILLC, University of Amsterdam; ILCC, University of Edinburgh(逻辑、语言与计算研究所,阿姆斯特丹大学; 信息实验室,爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在提高无标签大语言模型推理准确性,提出CANON方法将共识转化为token级监督,通过采样多个解决方案并以达到多数答案的方案为条件设置模型快照来监督。实验表明该方法大幅提升性能,还能迁移,改进非单纯分布锐化。

AI 中文摘要

采样多个解决方案并返回多数答案是提高无标签大语言模型推理准确性的最可靠方法之一,越来越多的方法将这种共识信号转化为训练监督。然而,现有方法仅以受限形式使用共识。我们提出了CANON,一种无标签训练方法,将共识转化为密集的、token级监督。对于每个无标签提示,CANON采样多个解决方案,提取多数答案,并以达到该答案的解决方案为条件设置模型的冻结快照;然后这个基于共识的教师在每个token上监督模型自身的展开。在数学和科学推理基准上的实验表明,CANON将pass@1提高了多达12分,在计算量仅为无标签强化学习七分之一的情况下比其高出6分,接近基于正确答案训练的教师模型;在聚合无标签数据上训练后,它可迁移到保留基准上匹配使用正确标签的训练方法。分析表明这些改进并非纯粹的分布锐化:训练后,模型能解决之前32次尝试中都未解决的问题,且其多数投票本身也变得更准确。

英文摘要

Sampling multiple solutions and returning the majority answer is among the most reliable ways to improve the reasoning accuracy of large language models without labels, and a growing family of methods converts this consensus signal into training supervision. However, existing approaches use consensus only in restricted forms: as a filter that selects solutions for fine-tuning, as a preference between answers, or as a scalar reward for reinforcement learning, discarding most of the information that the agreeing solutions contain. We present CANON (Consensus-ANchored self-distillatiON), a label-free training method that turns consensus into dense, token-level supervision. For each unlabeled prompt, CANON samples multiple solutions, extracts the majority answer, and conditions a frozen snapshot of the model on a solution that reaches it; this consensus-anchored teacher then supervises the model on its own rollouts at every token. Experiments on mathematical and scientific reasoning benchmarks show that CANON improves pass@1 by up to 12 points, outperforming label-free reinforcement learning by 6 points at a seventh of its compute and approaching a teacher conditioned on gold solutions; trained on pooled unlabeled data, it transfers to held-out benchmarks, matching training methods that use gold labels. Analysis suggests that the improvements are not pure distribution sharpening: after training, the model solves problems it previously never solved in 32 attempts, and its majority vote itself becomes more accurate.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑