AI 中文总结
该研究发现语言模型存在校准错误读出与行为隐藏知识,内部线性探测能读出正确逻辑结论,而模型输出因阈值偏移失效,修正参数可提升准确率,还区分出三种知识状态。
AI 中文摘要
一个0.6B的语言模型被要求验证1200个逻辑结论(其中一半是有效的,另一半因单个语义编辑而被破坏),它每次都回答“是”。从行为来看,它无法区分任何内容;而对其隐藏状态进行的线性探测以0.96的AUC读出了正确的判断,且该结果可迁移至未见过的逻辑结构,并能区分由真实结论的完全相同词汇构建的干扰项(AUC为0.90)。我们探究了判断在何处丢失,发现主要的失败原因是一个单一标量。该判断沿着一条对齐良好的读出方向保留在模型自身的输出logits中(边际AUC为0.89);一个饱和的决策阈值,偏移了+4.6个标准差,将其抹去。该诊断具有可推广性:在一个包含5个模型、3个系列的析因实验的90种语义标签配置中,行为准确率崩溃为阈值偏移的单一函数(斯皮尔曼相关系数为-0.93),而边际排名的变化则小得多。在13倍的规模范围内,内部知识达到饱和,而自由形式的行为则呈非单调变化:一个8B模型因答案通道故障而非阈值问题,其表现不如其4B的同类模型;强制选择准确率则呈单调变化。该诊断是可操作的:一个从未在评估结构上拟合的单参数修正,将行为从50%提升至81%(针对0.6B模型);校准后的边际解码在8B模型上恢复了94%的准确率;少样本提示的作用方式相同,将阈值从+4.6个标准差重新调整至0.0个标准差,同时保留排名。将探测与边际进行比较,可区分出三种状态:隐藏型、校准错误型和未检测型。在一个构建得使干扰项不带有表面线索的迷宫任务中,该审计正确报告了第三种状态。在标准生成设置中,答案表面特征和启发式标签在无需任何内部访问的情况下,即可复现已发表的探测结果。
英文摘要
A 0.6B language model, asked to verify 1,200 logical conclusions (half valid, half corrupted by a single semantic edit), answers YES every time. Judged by behavior it discriminates nothing; linear probes on its hidden states read the correct verdict at 0.96 AUC, transferring to unseen logical structures and separating foils built from exactly the words of the true conclusion (0.90). We ask where the verdict is lost, and find the dominant failure is a single scalar. The verdict survives to the model's own output logits (margin AUC 0.89) along a well-aligned readout direction; a saturated decision threshold, offset by +4.6 sigma, erases it. The diagnosis generalizes: across 90 semantic-label configurations of a five-model, three-family factorial, behavioral accuracy collapses onto a single function of threshold offset (Spearman -0.93) while margin ranking moves far less. Across a 13x scale range, internal knowledge saturates while free-form behavior is non-monotone: an 8B model underperforms its 4B sibling through an answer-channel failure rather than the threshold; forced-choice accuracy is monotone. The diagnosis is actionable: a one-parameter correction, never fit on evaluated structures, repairs behavior from 50% to 81% (0.6B); calibrated margin decoding recovers 94% at 8B; few-shot prompting works the same way, recentering the threshold (+4.6 sigma to 0.0 sigma) while preserving ranking. Comparing probe to margin separates three regimes: concealed, miscalibrated, and undetected. On a maze task built so foils carry no surface cues, the audit correctly reports the third. In the standard generation setting, answer-surface features and heuristic labels reproduce published probing results without any internal access.