arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

答案并非论证

The Answer Is Not the Argument

Will Yeadon, Sergio Juárez, Paul Mackay, T. J. Dowling, Elise Agra, Oto-obong Inyang, Arin Mizouri, Craig P. Testrow

arXiv 2609.00264首次发表:更新:

发表机构

Durham University; University of Vigo(杜伦大学; 维戈大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究探究答案访问对AI思维链监控的影响,发现其提升结论一致性检查能力,但会低估含真实错误的正确答案轨迹的错误,提示可信答案评估或高估监控能力。

AI 中文摘要

针对AI监督,思维链监控被提出,但评估常为监控器提供可信的参考答案。我们探究答案访问是提升了推理验证能力,还是主要暴露了错误结论。我们从三个前沿模型收集了79道Humanity's Last Exam物理题的237个带步骤编号的解答,未插入错误,且独立标注了最终答案正确性与首个错误步骤。参考标准结合了物理学家标注、独立LLM辩论及隐藏来源裁决,由此得到24条关键轨迹,其中答案正确但轨迹包含真实错误。8个LLM监控器对轨迹进行盲测,使用未验证或经认证的答案,或在盲承诺后操作。认证使平均平衡准确率从0.637提升至0.796,首个错误定位准确率从0.261升至0.379;在错误答案轨迹上,认证使召回率(标记为错误的错误轨迹比例)从0.653升至0.951,但在关键轨迹上从0.521降至0.438,所有8个监控器均呈现该对比方向(自助抽样95%置信区间[+0.256, +0.506])。盲承诺后,展示答案的监控器将93.8%先前通过的错误答案轨迹标记为错误,但仅标记18.0%的关键轨迹。因此,答案访问提升了结论一致性检查,而非对支撑论证的独立验证。对于AI安全,这些轨迹提供了奖励黑客行为的良性类比:可接受的输出无法证明产生该输出的过程是合理的。尽管此处研究的错误是普通的、大多非关键的而非对抗性的,但当可接受输出掩盖了不合理推理时,可信答案评估可能同样高估监控能力。

英文摘要

Chain-of-thought monitoring is proposed for AI oversight, yet evaluations often provide monitors with a trusted reference answer. We ask whether answer access improves verification of the reasoning or mainly supplies information about its conclusion. We collected 237 naturally generated, step-numbered solutions to 79 Humanity's Last Exam physics questions and independently labelled final-answer correctness and the first false step. Eight LLM monitors evaluated the traces with varying access to the reference answer. Certification raised mean balanced accuracy from 0.637 to 0.796, but its effect on error detection depended strongly on the conclusion: recall increased by +0.299 on wrong-answer traces, while there was no evidence of improvement on correct-answer traces containing a reasoning error (-0.083, 95% CI [-0.196, +0.030]). We then held the reasoning trace fixed in a seven-monitor certificate-congruence intervention. Replacing the true certificate with the trace's own incorrect conclusion reduced flagging by 0.659 (95% CI [0.602, 0.711]) and left flagging 0.389 below the answer-blind level. Conversely, a conflicting false certificate increased flagging of clean traces by 0.580, with 82.9% of newly flagged cases assigning the alleged error to an interior reasoning step. Trusted-answer access can therefore make monitoring appear substantially stronger because aggregate performance combines independent reasoning verification with a powerful certificate-conclusion consistency signal.

Comments25 pages, 12 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑