更多正确质量,更差答案:幂采样为何会失效及如何修复
More Correct Mass, Worse Answers: Why Power Sampling Can Fail and How to Fix It
浏览论文内容
中文总结 AI 辅助
研究发现幂采样存在悖论,即提升正确轨迹概率却降低下游推理性能,归因于剂量与覆盖不匹配,提出修复后的支持保留幂采样器可逆转损失并优于标准多采样推理。
中文摘要 AI 辅助
幂采样可增强语言模型在完整生成轨迹上的分布,提供一种无需验证器即可在推理时提升推理能力的方法,也有望作为通用前端服务于广泛的下游采样方法。然而,我们发现一个显著悖论:幂采样可将更多概率质量推向正确轨迹,却会降低其本应增强的下游推理性能。以自一致性为代表案例,我们观察到在不同模型和推理基准上,准确率下降幅度最高达18.5个百分点。我们将此悖论归因于两类不匹配:剂量不匹配源于固定指数会在不同问题上引发截然不同的分布变化幅度;覆盖不匹配源于全局增强会将质量集中在少数主导路径上:因此,常被视为多样性保留证据的高pass@k值,可能与下游聚合、搜索及选择所需的广泛推理路径支持的丧失共存。基于此诊断,我们用变形控制、支持保留的幂目标取代均匀轨迹指数化,该目标可校准不同问题的增强程度,同时限制对中等概率路径的抑制。在与加权自一致性相同预算的实例中,修复后的采样器逆转了全局幂采样导致的损失,且在推理基准上优于标准多采样推理。
英文摘要
Power Sampling sharpens a language model's distribution over complete generation trajectories, offering a verifier-free way to improve reasoning at inference time. It also has the potential to serve as a general-purpose front end for a broad range of downstream sampling methods. However, we uncover a striking paradox: Power Sampling can drive more probability mass toward correct trajectories while degrading the downstream inference it is intended to enhance. Using self-consistency as a representative case, we observe accuracy drops of up to 18.5 percentage points across models and reasoning benchmarks. We trace this paradox to two mismatches. Dose mismatch arises because a fixed exponent induces drastically different amounts of distributional change across problems. Coverage mismatch arises because global sharpening concentrates mass on a narrow set of dominant paths: high pass@k, often interpreted as evidence of preserved diversity, can therefore coexist with the loss of broad reasoning-path support required for downstream aggregation, search, and selection. Guided by this diagnosis, we replace uniform trajectory exponentiation with a deformation-controlled, support-preserving Power target that calibrates sharpening across problems while limiting the suppression of moderate-probability paths. In a same-budget instantiation with weighted self-consistency, the repaired sampler reverses the losses caused by global Power and outperforms standard multi-sample inference across reasoning benchmarks.
发表机构
- State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能 State Key Laboratory)
机构由 AI 辅助整理,请以论文原文为准。