发表机构
University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有SMC等权重重采样会剪除低权重轨迹、降低推理路径多样性的问题,提出Chopthin-共识功率采样(CCPS),通过限制权重比并保留不等权重来保持多样性,结合语义多数选择提升最终答案准确率,在多个基准上显著优于基线。
AI 中文摘要
通过序贯蒙特卡洛(SMC)进行推理时功率采样可以显著提升大语言模型(LLM)的推理能力,而无需进行后训练。然而,许多现有的SMC方法依赖于等权重重采样,这可能会激进地剪除低权重轨迹,丢弃潜在正确的推理路径,并降低搜索空间的谱系多样性。为解决这一问题,我们引入了Chopthin-共识功率采样(CCPS)。我们的方法将Chopthin重采样器应用于LLM解码:它不是均衡权重并强制不必要的粒子复制,而是对最大和最小权重之间的比率设置上限,并向前传递不等的权重。这种有针对性的干预保留了更丰富的不同推理路径集合,在条件期望下保持加权SMC近似不变,并保证重采样后有效样本量(ESS)的下界。为了充分利用这种丰富的种群,我们采用了一种语义多数选择机制,该机制合并令牌相同的最终轨迹,对语义等价的答案进行聚类,并返回由最多不同轨迹支持的答案。在三个开放权重模型和五个推理基准上的评估表明,Chopthin在15个设置中的13个中提高了oracle覆盖率。结合语义多数选择,CCPS在15个设置中的14个中达到或超过了Power-SMC基线的最终答案准确率,绝对提升高达10.6个百分点。这些发现表明,保持多样性的重采样和多样性感知的选择是互补的机制,用于无需训练的LLM推理。代码可在以下网址获取:此HTTP URL。
英文摘要
Inference-time power sampling via Sequential Monte Carlo (SMC) can substantially improve large language model (LLM) reasoning without requiring post-training. However, many existing SMC approaches rely on equal-weight resampling, which can aggressively prune low-weight trajectories, discarding potentially correct reasoning paths and degrading the genealogical diversity of the search space. To address this, we introduce Chopthin-Consensus Power Sampling (CCPS). Our method applies the Chopthin resampler to LLM decoding: rather than equalizing weights and forcing unnecessary particle duplication, it enforces an upper bound on the ratio between the largest and smallest weights and carries the unequal weights forward. This targeted intervention preserves a richer set of distinct reasoning paths, keeps the weighted SMC approximation unchanged in conditional expectation, and guarantees a lower bound on the post-resampling effective sample size (ESS). To fully exploit this enriched population, we employ a semantic-majority selection mechanism that merges token-identical final trajectories, clusters semantically equivalent answers, and returns the answer supported by the largest number of distinct trajectories. Evaluating across three open-weight models and five reasoning benchmarks, we show that Chopthin increases oracle coverage in 13 of 15 settings. Combined with semantic-majority selection, CCPS matches or exceeds the final-answer accuracy of the Power-SMC baseline in 14 of 15 settings, delivering absolute gains of up to 10.6 percentage points. These findings demonstrate that diversity-preserving resampling and diversity-aware selection are complementary mechanisms for training-free LLM reasoning. Code is available at github.com/MinooAhmadii/chopthin-consensus-power-sampling.
CommentsAccepted at the COLM 2026 Workshop on Efficient Reasoning