发表机构
Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过判决替换干预在自一致性面板中分析验证者投票的符号枢轴值,发现不同模型验证者产生正增益而角色反转产生负增益,并利用精确分解解释符号,揭示有害判决替换而非普遍有害的多数投票。
AI 中文摘要
替换一张选票只有在那些由单票决定的查询上才能改变多数决策;这一结构性事实不需要独立性假设。我们使用一种标记的、判决式干预来研究这种变化的符号:在$k{=}7$的自一致性面板中,一个正确性信号替换一个正确性指示选票。这种诊断性干预不同于部署中的答案身份多数投票。一个主要的MATH-500实验($n{=}570$)中,不同模型的验证者获得了$+24.2$百分点的枢轴增益,而角色反转的配置则给出$-11.2$个百分点;一个探索性的代码压力测试(9个任务中的14个枢轴行)给出$-24.5$个百分点。一个精确的符号增益分解通过验证者在特定状态下的准确性和两个单票计票状态的组成来解释所有观察到的符号,而不是全局准确性或模型来源。同源信号在枢轴层上的准确性下降(主要配置中从65%降至44%),而错误相关性提供了一个描述性的错误关联诊断。受控退化以及$k\in\{3,5,7\}$的子集敏感性分析探究了该模式在该核算下的稳定性。在评估的平局错误答案身份多数分析中,结构性零和强验证者收益持续存在,但角色反转的损害减弱至$-0.9$个百分点且不显著。因此,结果确立了有害的判决替换,而非普遍有害的部署多数投票,并激发了一个可测试但未验证的关于负过程奖励模型权重的假设。
英文摘要
Replacing one ballot can change a majority decision only on queries decided by a single vote; this structural fact requires no independence assumption. We study the sign of that change using a labeled, verdict-style intervention: one correctness signal replaces one correctness-indicator ballot in $k{=}7$ self-consistency panels. This diagnostic intervention is not identical to deployed answer-identity plurality. A primary MATH-500 experiment ($n{=}570$) gives a different-model verifier a $+24.2$pp pivotal gain, whereas a role-reversed configuration gives $-11.2$pp; an exploratory code stress test (14 pivotal rows across 9 tasks) gives $-24.5$pp. An exact signed-gain decomposition accounts for all observed signs through the verifier's state-specific accuracy and the composition of the two one-vote tally states, rather than global accuracy or model provenance. Same-source signals lose accuracy on the pivotal stratum (65$\to$44\% in the primary configuration), while error correlations provide a descriptive error-association diagnostic. Controlled degradation and a $k\in\{3,5,7\}$ subset sensitivity analysis probe the stability of the observed pattern around this accounting. Under the evaluated ties-incorrect answer-identity plurality analysis, the structural zero and strong-verifier benefit persist, but the role-reversed harm attenuates to $-0.9$pp and is not significant. The results therefore establish harmful verdict substitution, not universally harmful deployed plurality, and motivate a testable but unverified hypothesis for negative process-reward-model weights.