通过优化引导模型缓解数学推理中PRM引导搜索的过优化问题
Mitigating Over-Optimization in PRM-Guided Search in Mathematical Reasoning by Optimizing the Guide
- Northwestern University(西北大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对PRM引导数学推理搜索的过优化问题,提出极大极小PRM引导搜索方法,无需微调即可提升搜索性能17%-35%,在多数设置中优于基线方法。
AI中文摘要:
过程奖励模型(PRMs)为基于搜索的推理提供密集的步骤级引导,可将推理时的计算资源分配给有前景的部分解。但近期研究表明,PRM引导搜索会对不完美的过程奖励过优化,修剪可行轨迹的同时扩展虚假轨迹。本研究从理论上证明,直接利用PRM评分易受验证器噪声影响,存在极值效应:随着推理深度增加,不可行前缀更可能获得虚假高分。因此,我们将PRM引导搜索表述为关于合理奖励扰动的鲁棒优化问题,称为极大极小PRM引导搜索,形成一种无需微调的鲁棒过程监督方法,在步骤级评分存在噪声时保留有前景的备选方案。极大极小PRM引导搜索通过降低对过优化PRM异常值的敏感性缓解该失效模式。无需微调或在线适配,极大极小搜索平均可将PRM引导搜索提升17%-35%,在16个设置中的14个里优于结果级和步骤级基线。我们的源代码可在该https URL获取。
英文摘要:
Process reward models (PRMs) provide dense step-level guidance for search-based reasoning, enabling inference-time compute to be allocated toward promising partial solutions. However, recent evidence suggests that PRM-guided search can over-optimize imperfect process rewards, pruning viable trajectories while expanding spurious ones. In this work, we theoretically show that directly leveraging PRM score is vulnerable to verifier noise through an extreme-value effect: non-viable prefixes become more likely to receive spuriously high scores as reasoning depth increase. Therefore, we formulate the PRM-guided search as a robust optimization problem over plausible reward perturbations, termed maximin PRM-guided search, leading to a training-free robust process supervision method that preserves promising alternatives when step-level scores are noisy. Maximin PRM-guided search mitigates this failure mode by reducing sensitivity to over-optimized PRM outliers. Without fine-tuning or online adaptation, maximin search consistently improves the PRM-guided search by 17-35\% on average, outperforming outcome- and step-level baselines in 14 out of 16 settings. Our source code is available at https://github.com/tjoo512/maximin-search.