发表机构
Efficient Computation Inc.(高效计算公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出用线性探针检测并修复语言模型在上下文绑定任务中的错误,通过探针输出不一致分数提升失败检测,并利用残差流干预提高准确率。
AI 中文摘要
错误的答案无法表明模型是缺乏所需信息,还是拥有信息却未能使用。在一个实体-义务绑定任务中,语言模型可能输出错误的提示提供的绑定,而线性探针可以从其冻结的隐藏状态中恢复正确的绑定。我们测量了这种情况在16个公开检查点中发生的频率,每个检查点使用三个随机种子进行评估。我们在训练集上拟合探针,在验证集上选择其层,并在独立的测试集上报告结果。在每个模型出错的试验中,探针准确率超过了严格存在义务基线1/K = 0.125,高出+0.196(95%置信区间[+0.101, +0.296],基于模型进行自助抽样)。查询实体反事实排除了标记存在和近期性的影响。基于探针输出不一致符号构建的分数,在失败检测上比模型自身置信度高出+0.079 AUROC(95%置信区间[+0.036, +0.126])。原始探针置信度相对于模型置信度没有可测量的改进。在没有金标签的情况下,将残差流导向探针解码的绑定,在测试的所有八个模型上平均准确率提高了+0.168(95%置信区间[+0.066, +0.280])。最近的研究报告称探针检测到的错误对干预具有抵抗力,而我们发现上下文绑定是探针可操作的场景。
英文摘要
A wrong answer does not show whether the model lacked the needed information or held it and failed to use it. On an entity-obligation binding task, a language model can emit an incorrect prompt-supplied binding while a linear probe can recover the correct one from its frozen hidden state. We measure how often this occurs across 16 public checkpoints, each evaluated with three seeds. We fit a probe on a training fold, select its layer on a validation fold, and report results on a disjoint test fold. On the trials each model gets wrong, probe accuracy exceeds the strict present-obligation baseline, 1/K = 0.125, by +0.196 (95% CI [+0.101, +0.296], bootstrapped over models). A query-entity counterfactual rules out token presence and recency. A score built from the sign of probe-output disagreement improves failure detection over the model's own confidence by +0.079 AUROC (95% CI [+0.036, +0.126]). Raw probe confidence gives no measurable improvement over model confidence. Steering the residual stream toward the probe-decoded binding, with no gold label, raises accuracy on all eight models tested by a mean of +0.168 (95% CI [+0.066, +0.280]). Where recent studies report that probe-detected errors are resistant to interventions, we find that in-context binding is a setting in which probes are actionable.