arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

表征推理时PRM剪枝片段嫁接失效的一种配置:来自三个推理语言模型的证据

Characterizing a Configuration Where Inference-Time PRM-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMs

Khawaja Murad ul Hassan, Mehran Ebrahimi

arXiv 2610.00047首次发表:更新:

发表机构

Ontario Tech University(安大略理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过等价性检验证明,在特定操作点下,PRM剪枝片段嫁接(PPFG)对并行思维链推理无增益,并分析其失效原因,贡献了推理时机制零效应的检验模板。

AI 中文摘要

并行思维链中的多样性坍缩促使了基于自然设计的推理时干预:当过程奖励模型(PRM)剪枝一条链时,其高PRM前缀被提取并逐字嫁接为上下文演示,进入仍在解码的兄弟链。我们将这一机制——PRM剪枝片段嫁接(PPFG)——隔离为跨轨迹步骤级转移的最成本最小化操作化实现,并在先前片段嫁接工作仅在额外补偿成分下报告收益的操作点进行测试。在Qwen2.5-7B-Instruct与Math-Shepherd在完整MATH500(n=500,三个随机种子)上,PPFG的停滞定向和随机定向变体在每项测量轴上与独立并行思维链基线在统计上无显著差异。我们表征其原因:对322次停滞规则注入事件的四桶分类显示,仅14%针对真正挣扎的链;其余落在已成功、接近完成或处于平坦PRM平台的链上,这些状态是救援嫁接无法改变的。没有复合门控细化能同时实现良好定向触发和足够密度,随机对照在2.4倍触发率下匹配相同平价,因此惰性并非启发式特定。该发现复制于三个基础语言模型、六个基准、第二个PRM以及兼容性门控扫描;双单侧检验分析将平价提升为所有十二个Qwen/LLaMA单元上的正等价。逐事件抽查发现注入链以2.75倍匹配步骤率被剪枝,但幸存兄弟反事实发现无总体水平补偿。事后预言机将选择PPFG而非独立方法带来的每问题增益上界限定为+0.13个百分点。我们贡献了一个等价性检验模板,用于建立推理时机制零效应,每项主张均限定于其测试操作点。

英文摘要

Diversity collapse in parallel chain-of-thought has motivated inference-time interventions built on a natural design: when a process reward model (PRM) prunes a chain, its high-PRM prefix is extracted and grafted verbatim as an in-context demonstration into a still-decoding sibling. We isolate this mechanism, PRM-Pruned Fragment Grafting (PPFG), as the most cost-minimal operationalization of cross-trajectory step-level transfer, and test it at the operating point where prior fragment-grafting work reports gains only under additional compensating ingredients. On Qwen2.5-7B-Instruct with Math-Shepherd on full MATH500 (n=500, three seeds), PPFG in both stagnation- and random-targeting variants is statistically indistinguishable from an independent parallel-CoT baseline on every measured axis. We characterize why: a four-bucket classification of 322 stagnation-rule injection events shows only 14% targeted a genuinely struggling chain; the rest landed on chains that had already succeeded, were near completion, or sat on a flat PRM plateau, states a rescue graft cannot change. No compound-gate refinement jointly achieves well-targeted firing and adequate density, and a random control matches the same parity at 2.4x the firing rate, so the inertness is not heuristic-specific. The finding replicates across three base LMs, six benchmarks, a second PRM, and a compatibility-gate sweep; two-one-sided-tests analysis promotes the parity to positive equivalence on all twelve Qwen/LLaMA cells. A per-event spot-check finds injected chains prune at 2.75x the matched-step rate, but a surviving-sibling counterfactual finds no population-level compensation. A hindsight oracle bounds any per-problem gain from choosing PPFG over independent at +0.13 pp. We contribute an equivalence-testing template for establishing inference-time mechanism nulls, with every claim scoped to its tested operating point.

Comments24 pages, 4 figures, 22 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑