知-言差距:当探测模型发现错误而置信度未察觉时
The Knowing-Saying Gap: When Probes See Errors that Confidence Misses
浏览论文内容
中文总结 AI 辅助
该研究发现线性探测模型能精准检测语言模型的上下文破坏,但无法可靠预测最终答案正确性,推翻了相关假设,表明基于探测模型的监控需结合模型与错误类型的路由策略。
中文摘要 AI 辅助
线性探测模型能以近乎完美的准确率检测出语言模型中被破坏的上下文,但这无法转化为可靠的故障预测,由此产生的分离对部署监控具有直接影响。在多跳算术链任务中,检测到破坏的探测模型对最终答案正确性并无指示作用;被强制采用结构化置信度格式的模型会崩溃为两个值,其错误率难以区分;探测模型在各跳间的持续性无法区分正确与错误结果,这推翻了我们预先注册的“持续性优于峰值”假设。这种“知晓但不表达”的模式在包括推理模型在内的多个模型家族中普遍存在。作为实时监控工具,基于探测模型的干预措施具有明显的模型和错误类型依赖性:分支选择(branch-and-pick)在所有模型中均为净正向效果,且是唯一对Llama-3.1-8B无破坏的干预方式(拯救4个错误答案,未破坏任何正确答案);而重新提示(reprompt)和替换先前内容(replace-prior)破坏正确答案的概率与拯救错误答案的概率大致相当。基于探测模型的监控是对口头表达置信度的必要补充,但没有单一干预措施占主导,可部署的解决方案是基于模型和错误类型的路由策略。
英文摘要
Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The result is a dissociation with direct implications for deployment monitoring. Across multi-hop arithmetic chains, probes that detect corruption turn out to be uninformative about final answer correctness; models forced into structured confidence formats collapse to two values with indistinguishable error rates; and probe persistence across hops fails to separate correct from incorrect outcomes, refuting our pre-registered "persistence beats peak" hypothesis. This pattern of knowing but not saying generalises across model families including reasoning models. As a real-time monitor, probe-based interventions are sharply model and error-type dependent: branch-and-pick is net-positive across models and uniquely non-breaking on Llama-3.1-8B (4 rescued, 0 broken), while reprompt and replace-prior break correct traces at roughly the rate they rescue wrong ones. Probe-based monitoring is a necessary complement to verbalised confidence, but no single intervention dominates, and the deployable answer is model-aware, error-type-aware routing.