arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16799cs.CLcs.LG

自我判断混淆下的正确性探测诊断

Diagnosing Correctness Probes under Self-Judgement Confounding

Yi-Long Lu

首次发表
浏览论文内容

中文总结 AI 辅助

研究语言模型中隐藏状态读数预测正确性时,因客观正确性与模型自我判断混淆致信号模糊。构建冲突案例,估计相关方向极性,发现与SJ相关方向跨域转移高于机会水平,OC相关方向则相反,且转移不对称,表明仅可转移性不能确立客观正确性语义。

中文摘要 AI 辅助

隐藏状态读数可预测语言模型输出是否正确,但客观正确性(OC)通常与模型自身的自我判断(SJ)一致,导致解码信号语义模糊。我们构建了冲突案例,其中OC和SJ预测相反的读数顺序。在高置信度分歧中,传统的正确性标记对比通常将错误/自我认可的回答排在正确/自我拒绝的回答之上,遵循SJ而非OC。我们估计了与SJ和OC相关的阶乘方向,并在数学推理和事实回忆中评估它们的极性。在四个高达14B参数的指令调整模型中,与SJ相关的方向在每个模型的两个跨域方向上的转移高于机会水平,而与OC相关的方向在每个相应条件下对预期OC排序的点估计低于机会水平。这种转移不对称在中到后期层发展,在答案可能性、序列长度和空方向控制下持续存在,并扩展到MMLU和二元TruthfulQA,无需目标域方向拟合。在研究的模型和诊断子集中,最可靠可转移的组件保留了与SJ相关的极性。因此,仅可转移性并不能确立客观正确性语义。

英文摘要

Hidden-state readouts can predict whether language-model outputs are correct, but objective correctness (OC) usually agrees with the model's own self-judgement (SJ), leaving the decoded signal semantically ambiguous. We construct conflict cases in which OC and SJ predict opposite readout orderings. On high-confidence disagreements, conventional correctness-labelled contrasts often rank incorrect/self-endorsed responses above correct/self-rejected responses, following SJ rather than OC. We estimate factorial SJ- and OC-associated directions and evaluate their polarity across mathematical reasoning and factual recall. Across four instruction-tuned models up to 14B parameters, the SJ-associated direction transfers above chance in both cross-domain directions for every model, whereas the OC-associated direction has a below-chance point estimate for the expected OC ordering in every corresponding condition. This transfer asymmetry develops across middle-to-late layers, persists under answer-likelihood, sequence-length, and null-direction controls, and extends to MMLU and binary TruthfulQA without target-domain direction fitting. Across the studied models and diagnostic subsets, the most reliably transferable component preserves SJ-associated polarity. Transferability alone therefore does not establish objective-correctness semantics.

↑