发表机构
National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究在冻结大语言模型中定位答案正确性信号,发现其集中于答案跨度且多信号族互补,融合可提升分布外鲁棒性,并用于门控检索控制器。
AI 中文摘要
语言模型会暴露内部信号,这些信号能够预测答案是否正确,并且可以在冻结模型的单次前向传播中读取,无需额外的生成过程。然而,现有的探针常常局限于单一信号族或单一层,并且在分布偏移下可能变得脆弱;在检索增强的场景中,许多专门的检测器转而针对段落忠实度,当检索到的证据无用或相互冲突时,这可能会与正确性产生分歧。因此,我们探究答案正确性在何处可读、哪些内部信号族承载了它,以及应如何组合这些信号。我们搜索了隐藏状态、令牌概率、残差流特征、注意力及其融合,并将选定的读出视为一种预测性测量,而非机制性定位。我们在闭卷和带上下文两种设置下分别进行此分析,因为上下文可能改变哪些读出具有信息量。一个一致的解剖结构浮现出来:正确性集中于答案跨度,即使在检索情况下也能从答案令牌中恢复,并且各信号族互补地承载它,因此融合它们在最有助于分布外场景,此时单一信号最弱。该协议在两个骨干网络上有效,并作为下游用途之一门控检索控制器。
英文摘要
Language models expose internal signals that predict whether an answer is correct, readable from a single forward pass of a frozen model without additional generations. Yet existing probes often commit to one signal family or layer and can be brittle under distribution shift; in retrieval-augmented settings, many specialized detectors instead target passage faithfulness, which can diverge from correctness when retrieved evidence is unhelpful or conflicting. We therefore ask where answer correctness is readable, which internal signal families carry it, and how they should be combined. We search over hidden states, token probabilities, residual-stream features, attention, and their fusion, treating the selected readouts as a predictive measurement rather than a mechanistic localization. We run this analysis separately in closed-book and with-context settings, since context can change which readouts are informative. A consistent anatomy emerges: correctness concentrates in the answer span, recovered from the answer tokens even under retrieval, and the families carry it complementarily, so fusing them helps most out of distribution, where a single signal is weakest. The protocol is effective across two backbones and gates a retrieval controller as one downstream use.