无需说服的一致:单轮辩论抹去了验证所需的异议
Unanimity Without Persuasion: A Single Round of Debate Erases the Disagreement That Verification Needs
浏览论文内容
中文总结 AI 辅助
研究显示辩论小组无需说服即可达成一致,单轮辩论使一致率从39.5%跃升至95.2%但准确率几乎不变,抹去了验证所需的异议信号,建议在任何同行暴露前进行验证。
中文摘要 AI 辅助
辩论小组可以在不变得更正确的情况下达成一致。这对下游安全机制是危险的:被替换的验证选票只能改变微弱优势的投票,而更丰富的仲裁者会失去异议这一自然定位信号。我们表明,一轮辩论就能抹去该资源,且无需说服。通过追踪一个由7名法官组成的异质小组,在600个代码正确性候选上经历一轮盲评和三轮辩论,固定队列上的一致率从第1轮的39.5%跃升至95.2%(占总崩溃的93.1%),而准确率变动不到1个百分点,且96.3%的判决翻转跟随显示的同行多数。基于执行的验证选票在辩论前的2,037个候选替换实例中纠正了8个,但在后续每一轮中纠正为零;到第3轮,每个错误决定都是一致的,抹去了曾标记小组三分之二错误的异议。相同队列的对照组解释了原因:无同行重新考虑重现了79.6%的崩溃,无推理的真实标签重现了91.5%,而随机标签将翻转引向其显示的任何内容;完整辩论条件在仅标签基础上增加了4.3个百分点(聚类95%置信区间0.7--8.1)。单轮崩溃在另外两次真实运行和两个假标签种子中重现,在小组规模3--7下保持,并出现在MATH-500中。解析失败集中在有争议的候选上(p<0.001),使得流失不可忽略。设计含义是操作性的:在任何第二轮评估或同行暴露之前进行验证,且绝不将辩论后的一致视为可靠性的独立证据。
英文摘要
A debate panel can become unanimous without becoming more correct. This is dangerous for downstream safeguards: a substituted verification ballot can change only narrow-margin votes, while richer arbiters lose disagreement as a natural targeting signal. We show that one debate round can erase that resource without requiring persuasion. Tracking a heterogeneous 7-judge panel through a blind round and three debate rounds on 600 code-correctness candidates, unanimity on a fixed cohort jumps from 39.5\% to 95.2\% in round 1 (93.1\% of the total collapse), while accuracy moves by less than one point and 96.3\% of verdict flips follow the displayed peer majority. An execution-based verification ballot corrects 8 of 2,037 pre-debate candidate-substitution instances but changes zero in every later round; by round 3 every wrong decision is unanimous, erasing dissent that had flagged two-thirds of the panel's errors. Identical-cohort controls explain why: no-peer reconsideration reproduces 79.6\% of the collapse, real labels without reasoning reproduce 91.5\%, and random labels steer flips toward whatever they display; the full-debate condition adds 4.3 percentage points over labels only (clustered 95\% CI 0.7--8.1). The one-round collapse reproduces in two additional real runs and two fake-label seeds, remains under panel sizes 3--7, and appears in MATH-500. Parse failures concentrate on contested candidates ($p<0.001$), making attrition non-ignorable. The design implication is operational: verify before any second-pass evaluation or peer exposure, and never treat post-debate unanimity as independent evidence of reliability.
发表机构
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。