arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

流程奖励模型的质量-多样性压力测试:存档覆盖率能与不能证明什么

Quality-Diversity Stress Tests for Process Reward Models:What Archive Coverage Can and Cannot Certify

Ibne Farabi Shihab, Fariya Afrin

arXiv 2608.08008首次发表:更新:

发表机构

Iowa State University; Kalinga Institute of Industrial Technology(爱荷华州立大学; 卡林加工业技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对流程奖励模型(PRMs)的漏洞,提出基于MAP-Elites的质量-多样性压力测试方法,验证了存档覆盖率的理论边界,发现Qwen2.5-Math-PRM-7B存在聚合依赖漏洞,通过配对LoRA协议修复了漏洞

AI 中文摘要

流程奖励模型(PRMs)对中间推理步骤进行评分,广泛应用于搜索、排序和训练,但优化过程可能会利用这些学习到的代理,在提升奖励的同时将正确推理转为错误推理。我们将PRM压力测试表述为质量-多样性搜索问题,采用MAP-Elites算法,在每个行为空间区域中保留最严重的正确性翻转编辑,同时将搜索覆盖率与利用覆盖率分离。我们刻画了此类存档可证明的内容:有限单元修复界、已覆盖单元的尾部风险和平均残留严重程度,但仅靠已覆盖比例无法界定剩余最差单元;在利普希茨(Lipschitz)修复后损失和度量覆盖审计下,残留值受限于存档拟合误差加上利普希茨常数乘以覆盖半径。受控场景验证了该证明及仅靠比例无法提供最差情况保证的结论。在真实PRM上,搜索揭示了Qwen2.5-Math-PRM-7B存在依赖聚合方式的漏洞:在平均池化下,填充产生44个严格利用项,最大增益为0.294,而最小读出下仅1个利用项;匹配的句法控制分离了该机制,RLHFlow值头模型呈现相同定性效应,最大增益为0.005。预先声明的配对LoRA修复协议将利用率从0.148降至0.037至0.074,将最差攻击从0.333降至0.177至0.212,提升排序AUROC且不降低best-of-4准确率,将增益归因于对抗微调而非存档多样性,独立非配对复现确认了该结果(44至1,干净拆分最差增益0.0092,MATH-500 41至0,干净排序40/40)

英文摘要

Process reward models (PRMs) score intermediate reasoning steps and are widely used for search, ranking, and training, but optimization can exploit these learned proxies by increasing reward while turning correct reasoning into incorrect reasoning. We formulate PRM stress testing as a quality-diversity search problem using MAP-Elites, retaining the most severe correctness-flipping edit in each behavior-space region while separating search coverage from exploit coverage. We characterize what such archives certify: finite-cell repair bounds covered-cell tail risk and average residual severity but cannot bound the worst remaining cell from covered fraction alone; under Lipschitz post-repair loss and metric-cover auditing, the residual is bounded by archive fitting error plus the Lipschitz constant times the covering radius. A controlled landscape validates this certificate and the impossibility of any fraction-only worst-case guarantee. On real PRMs, the search reveals an aggregation-dependent vulnerability in Qwen2.5-Math-PRM-7B: padding yields 44 strict exploits with maximum gain 0.294 under mean pooling versus one exploit under minimum readout; a matched syntactic control isolates the mechanism, and an RLHFlow value-head model shows the same qualitative effect with maximum gain 0.005. A predeclared paired LoRA repair protocol reduces exploit rates from 0.148 to 0.037 to 0.074, lowers the worst attack from 0.333 to 0.177 to 0.212, improves ranking AUROC without degrading best-of-4 accuracy, attributes gains to adversarial fine-tuning rather than archive diversity, and is confirmed by independent unpaired replications (44 to 1, clean-split worst gain 0.0092, MATH-500 41 to 0, clean ranking 40/40).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑