arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

校准下限:格式修复可在中小规模模型中伪装成自我修正

The Calibration Floor: Format Repair Can Masquerade as Self-Correction at Small-to-Mid Scale

Mingguang Chen, Bo Qu, Licheng Wang

arXiv 2608.04355首次发表:更新:

发表机构

DeepGrounding; AlphaAvatar(深 grounding; 阿尔法化身)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现中小规模语言模型的自我修正准确率变化多由格式修复导致而非推理变化,因果实验证实格式效应占比更高,内容仅为自我修正测量的次要部分。

AI 中文摘要

语言模型自我修正后的准确率变化通常被解读为推理能力的变化。我们表明,这种解读在答案提取边界处可能失效,并通过因果方式而非仅观察方式测试该失效情况。在Qwen3.5(0.8B-9B)、Gemma-4-12B以及两个通过API获取的前沿模型(腾讯Hy3、英伟达Nemotron-3-Ultra-550B)的29个主要单元加一个前沿分支中,我们将始终存在的修正准确率变化分解为内容余量(两个答案均可解析)和格式恢复/损失余量(可解析性变化)。在12个具有意义的不可解析答案率的单元中,格式效应超过内容效应(Wilcoxon检验p值为1.7e-3)。为进行因果测试,我们通过语法约束解码强制已生成的推理过程,确保每个答案在结构上均可解析:在14个单元中,这一操作缩小了朴素总效应与内容余量估计之间中位数71%的差距,其中2个单元完全收敛,而在两个效应最大的单元中保留了残差而非忽略。聚类模型证实,下限规模(0.8B/2B)模型比能力规模模型具有高得多的内容级变化和损害概率(p<1e-7)。严格在Qwen3.5上复现引用的置信门控协议,未复现其报告的增益,且显示相同的近零内容余量。对更大规模模型的前沿检查显示,格式主导性随规模增大而增强:尽管总效应最高达+0.275,但所有5个单元的内容余量均恰好为零,不过该分支的统计效力较低。基于内容余量的校准下限标准揭示了一种挤压:下限规模单元有提升空间但信号不足,能力规模单元有信号但提升空间极小;仅1个单元勉强可行,密封保留集增益可忽略。内容仅占该领域所测量的自我修正的一小部分。我们发布了相关工具、代码及衍生结果。

英文摘要

Accuracy changes after language-model self-revision are usually interpreted as changes in reasoning. We show this can fail at the answer-extraction boundary, and test the failure causally rather than only observationally. Across Qwen3.5 (0.8B-9B), Gemma-4-12B, and two frontier models via API (Tencent Hy3, Nvidia Nemotron-3-Ultra-550B) in 29 primary cells plus a frontier arm, we decompose the always-revise accuracy shift into a content margin (both answers parseable) and format-recovery/loss margins (parseability changes). On 12 cells with meaningful unparseable-answer rates, format effects exceed content effects (Wilcoxon p=1.7e-3). To test this causally, we force already-generated reasoning through grammar-constrained decoding so every answer is parseable by construction: across 14 cells this closes a median 71% of the gap between the naive total effect and the content-margin estimate, with two cells converging exactly and a residual on the two largest-effect cells reported rather than dismissed. A clustered model confirms floor-scale (0.8B/2B) models have far higher odds of content-level change and harm than capable-scale models (p<1e-7). Replicating a cited confidence-gating protocol verbatim on Qwen3.5 does not reproduce its reported gain and shows the same near-zero content margin. A frontier check on much larger models shows format-dominance intensifying with scale: content margin is exactly zero in all 5 cells despite total effects up to +0.275, though this arm is lower-powered. The calibration-floor criterion on the content margin reveals a squeeze: floor-scale cells have headroom but insufficient signal, capable-scale cells have signal but little headroom; only one cell is marginally viable, with negligible sealed-holdout gain. Content is a minority share of what the field has measured as self-correction. We release the instrument, code, and derived results.

Comments36 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑