发表机构
Stevens Institute of Technology(史蒂文斯理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出无需训练的决策流采样(DF-Sample)框架,通过构建层次化推理树并执行全局轨迹评估,解锁基座模型中潜在的高质量推理路径,在GPQA上以45.6%准确率超越GRPO等基线。
AI 中文摘要
大语言模型推理中的一个核心问题是:强化学习(RL)究竟赋予了模型真正的新能力,还是仅仅重塑了推理过程中已有知识的表达方式。基于分布锐化假说——该假说认为强化学习将概率质量重新分配给基座模型中已潜伏的高奖励轨迹——我们提出疑问:能否在不进行昂贵的强化学习微调的情况下解锁这些潜在路径?我们提出了决策流采样(DF-Sample),一种无需训练、无需数据的推理时框架,它构建一棵层次化推理树,对终端节点进行质量评分,并反向传播效用值以指导每个中间分支决策。与采用纯局部逐步选择的传统采样策略不同,DF-Sample 在提交到某条路径之前执行显式的全局轨迹评估,从而恢复标准解码所忽略的高质量但低概率的推理链。在 GPQA 上,DF-Sample 达到 45.6% 的准确率,超过了幂采样(38.9%)和 GRPO(39.9%),表明无需训练的方法可以胜过经过训练的方法。在三个模型和四个基准上,DF-Sample 始终优于基线,表明预训练基座模型中存在大量潜在的推理能力。
英文摘要
A central question in LLM reasoning is whether reinforcement learning (RL) instills genuinely new capabilities or merely reshapes how existing knowledge is expressed during inference. Building on the distribution-sharpening hypothesis, which holds that RL reallocates probability mass toward high-reward trajectories already latent in base models, we ask: can we unlock those latent paths without costly RL fine-tuning? We present Decision-Flow Sampling (DF-Sample), a training-free, data-free inference-time framework that constructs a hierarchical reasoning tree, scores terminal nodes for quality, and back-propagates utilities to inform each intermediate branching decision. Unlike conventional sampling strategies that make purely local step-wise choices, DF-Sample performs explicit global trajectory evaluation before committing to a path, recovering high-quality but low-probability reasoning chains that standard decoding overlooks. On GPQA, DF-Sample achieves 45.6% accuracy, surpassing power sampling (38.9%) and GRPO (39.9%), showing that a training-free method can outperform a trained one. Across three models and four benchmarks, DF-Sample consistently outperforms baselines, indicating substantial latent reasoning potential in pretrained base models.