广泛探索,锐利推理:通过采样将小模型推向前沿
Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling
浏览论文内容
中文总结 AI 辅助
本文提出并行功率退火(PPT)方法,通过并行退火实现功率锐化采样,平衡探索与利用,显著提升小模型推理质量,超越RL后训练模型,性能接近前沿模型。
中文摘要 AI 辅助
功率锐化采样是一种推理时的替代方案,用于替代强化学习(RL)后训练,以增强大型语言模型(LLM)的推理能力。在基础模型下,高概率序列被放大,无需参数更新或外部奖励,从而避免了RL的昂贵优化和不规则泛化。然而,这种方法面临一个根本性的探索-利用权衡,因为强锐化限制了探索,使采样器陷入看似合理但错误的推理轨迹,而弱锐化则使答案分布变得分散。为了解决这一权衡,我们引入了并行功率退火(PPT),通过并行退火实例化功率锐化的LLM采样。在不同锐化级别上并行运行多个相互作用的副本,允许低功率副本探索多样化的推理轨迹,而高功率链则进一步利用锐化目标所偏好的更高似然响应。具体来说,我们针对推理时采样定制了PPT,通过缓解先前功率采样器中识别出的截断偏差,并研究了在有限内存和计算预算下的有效交换策略。大量实验表明,PPT显著改善了单链功率锐化采样,并优于RL后训练模型,产生更高质量的推理轨迹,甚至达到与前沿模型相当的性能。
英文摘要
Power-sharpened sampling is an inference-time alternative to reinforcement-learning (RL) post-training for enhancing reasoning in large language models (LLMs). High-probability sequences are amplified under the base model without parameter updates or external rewards, avoiding the costly optimization and jagged generalization of RL. However, this approach faces a fundamental exploration--exploitation trade-off, as % strong sharpening restricts exploration, trapping samplers in plausible but incorrect reasoning trajectories, whereas weak sharpening leaves the answer distribution diffuse. To resolve this trade-off, we introduce \textbf{Parallel Power Tempering (PPT)}, instantiating power-sharpened LLM sampling via parallel tempering. Running multiple \emph{interacting} replicas in parallel at different sharpening levels allows lower-power replicas to explore diverse reasoning trajectories and higher-power chains to further exploit higher-likelihood responses favored by the sharpened target. Specifically, we tailor \method{} to inference-time sampling by mitigating a truncation bias, identified in prior power samplers, and investigate effective swap strategies under finite memory and compute budgets. Extensive experimentation shows that \method{} substantially improves single-chain power-sharpened sampling and outperforms RL-post-trained models, producing higher-quality reasoning traces and even achieving performance comparable to frontier models.
发表机构
- Georgia Institute of Technology(佐治亚理工学院)
- University of Texas at El Paso(德克萨斯大学埃尔帕索分校)
- ML Research, Morgan Stanley(摩根士丹利机器学习研究部)
机构由 AI 辅助整理,请以论文原文为准。