思维链熵作为可靠性信号:一项预注册的复现研究
Chain-of-Thought Entropy as a Reliability Signal: A Preregistered Reproduction
浏览论文内容
中文总结 AI 辅助
本研究预注册复现了思维链熵形状预测正确性而总降幅不能的分离现象,在全规模基准上验证形状信号,并揭示大小信号依赖设置。
中文摘要 AI 辅助
这项实证研究是对Zhao在2026年报告的分离现象的独立复现。大型语言模型思维链熵轨迹的形状能预测最终答案是否正确,而其总熵降幅的大小则不能。该分离现象值得复现,因为大小半部分仅基于一个模型在一个随机种子下的300个问题运行,而形状半部分则在两个基准测试及第二个模型家族上以全规模报告。在OSF预注册于任何确认性运行之前,本次复现跨越完整的GSM8K和MATH-500基准测试集,使用四个开放权重模型,包括一个原始研究未测试的推理蒸馏模型。形状信号得到复现。大小信号随设置而分化。在锚定模型上,单调与非单调链之间的准确率差距在GSM8K上为+9.6个百分点,在MATH-500上为+27.5个百分点,而总熵降幅与正确性的秩相关系数在GSM8K上为-0.018,在MATH-500上为+0.414。在推理蒸馏模型上,形状信号的二元形式大约在每百条链中触发一次,太少而无法估计注册的对比,而分级违规计数在该模型上仍具有预测性。在探索性比较中,最终步熵单独在所有八个模型-基准组合中通过ROC面积优于二元形状标志,并在原始报告的风险覆盖面积上,在六个或七个组合中优于,取决于原始未声明的积分范围。该研究在七个文档化的协议差异下,以全测试集规模贡献了形状信号的复现,绘制了大小信号成立和失败的设置图,并测量了原始未报告的四个协议依赖项。
英文摘要
This empirical study is an independent reproduction of the dissociation Zhao reported in 2026. The shape of a large language model's chain-of-thought entropy trajectory predicts whether the final answer is correct, while the magnitude of its total entropy drop does not. The dissociation merits reproduction because the magnitude half rests on a single 300-problem run with one model at one seed, while the shape half was reported at full scale on both benchmarks and on a second model family. Registered at OSF before any confirmatory run, the reproduction crosses the complete GSM8K and MATH-500 benchmark test sets with four open-weight models including one reasoning-distilled model of a kind the original did not test. The shape signal replicates. The magnitude signal divides by setting. On the anchor model the accuracy gap between monotone and non-monotone chains is +9.6 percentage points on GSM8K and +27.5 on MATH-500, while the rank correlation of the total entropy drop with correctness is -0.018 on GSM8K and +0.414 on MATH-500. On the reasoning-distilled model the binary form of the shape signal fires on about one chain in a hundred, too few to estimate the registered contrast, while the graded violation count remains predictive there. In an exploratory comparison the final-step entropy alone outperforms the binary shape flag in all eight model-by-benchmark cells by ROC area, and in six or seven by the risk-coverage area the original reports, depending on an integration range the original does not state. The study contributes a reproduction of the shape signal at full test-set scale under seven documented protocol differences, a map of the settings where the magnitude signal holds and fails, and measurements of four protocol dependencies the original does not report.
发表机构
- AI for Altruism
机构由 AI 辅助整理,请以论文原文为准。