发表机构
The Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在提升无训练语言模型推理能力,提出深度熵引导采样(DEGS)方法,利用逐层熵坍缩作为质量信号,结合序列似然性与深度熵结构,在多个模型和基准测试中达到无训练最优精度,在域外和难题测试中表现出色。
AI 中文摘要
强化学习已成为提升大语言模型推理能力的主导范式,但它需要昂贵的训练、精心策划的数据和奖励信号。近期工作表明,在测试时从锐化的基础模型分布中采样可恢复大部分强化学习的增益,但现有方法仅依赖输出层似然性,忽略了Transformer的内部前向传播动态。我们引入了深度熵引导采样(DEGS),这是一种无训练的测试时方法,它利用逐层熵坍缩作为内在质量信号。我们观察到更强的推理器,包括强化学习后训练的变体,表现出独特的“后期坍缩”:在收敛之前,对数透镜解码熵在较深层之前一直保持升高。我们定义了一个序列坍缩深度$D(\mathbf{x})$和一个联合目标$\pi(\mathbf{x}) \propto p(\mathbf{x})^\alpha \exp(\beta D(\mathbf{x}))$,它将序列似然性与这种深度熵结构相结合,并在MCMC幂采样框架(DEGS-MCMC)中实例化。在三个开放权重模型和四个推理基准上,这种接近机会的每个候选信号在采样轨迹上复合,达到了无训练的最优精度,在域外和更难的分割上增益最大,而单独的似然性在此处不足,且仅需个位数百分比的挂钟开销。在GRPO训练的数学分割上,DEGS略落后于内部GRPO参考,但在所有三个模型的GPQA域外测试中超过了它,无需任何训练、奖励模型或标记数据。
英文摘要
Reinforcement learning (RL) has become the dominant paradigm for improving the reasoning capabilities of large language models, but it requires expensive training, curated data, and reward signals. Recent work shows that sampling from sharpened base-model distributions at test time recovers much of the RL gain, yet existing methods rely solely on output-layer likelihoods and ignore the transformer's internal forward-pass dynamics. We introduce Depth-Entropy Guided Sampling (DEGS), a training-free, test-time method that exploits layer-wise entropy collapse as an intrinsic quality signal. We observe that stronger reasoners -- including RL-posttrained variants -- exhibit a distinctive "late collapse": logit-lens decoded entropy stays elevated until deeper layers before converging. We define a per-sequence collapse depth $D(\mathbf{x})$ and a joint objective $π(\mathbf{x}) \propto p(\mathbf{x})^α\exp(βD(\mathbf{x}))$ that combines sequence likelihood with this depth-entropy structure, instantiated inside an MCMC power-sampling framework (DEGS-MCMC). Across three open-weight models and four reasoning benchmarks, this near-chance per-candidate signal compounds over the sampling trajectory into state-of-the-art training-free accuracy, with gains largest out of domain and on the harder splits -- exactly where likelihood alone falls short -- at single-digit-percent wall-clock overhead. DEGS narrowly trails an in-house GRPO reference on the math splits GRPO was trained for, yet surpasses it out of domain on GPQA for all three models, without any training, reward model, or labeled data.