发表机构
University of Houston(休斯顿大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究揭示语言模型隐藏状态中答案正确性编码为几何方向,通过计算位移均值得到评分方向,无需参数更新即可提升事实基准和TruthfulQA上的评分性能,并发现不同正确性概念在表示空间中正交,解释校准失败源于路由问题。
AI 中文摘要
答案正确性在语言模型的隐藏状态中被编码为可恢复的几何方向。我们展示了,在不进行参数更新的情况下,从五十个带标签的样本中,在约70%模型深度处计算从错误答案表示到正确答案表示的位移均值,可得到一个评分方向,该方向在事实基准(ARC-Challenge和MMLU)上比零样本对数概率评分高出最多+32.0个百分点,在TruthfulQA上高出+38.1至+51.8个百分点,涵盖三个架构家族(Llama、Qwen、Gemma)中参数规模从1B到8B的五个模型。该方法每个候选答案仅需一次前向传播和一次点积,推理时不进行生成。作为幻觉检测器应用于单个(问题,答案)对时,恢复的方向达到0.693的AUROC,而对数概率评分为0.578。我们进一步发现,事实推理、领域知识和校准的真实性的正确性方向在表示空间中接近正交,揭示语言模型将几何上独立的子空间分配给质量上不同的正确答案概念,且分离程度随架构而变化。这一结构解释了观察到的迁移模式——在事实问题上校准的方向在任务类型内迁移,但不会跨类型迁移——并表明LLM校准失败可能反映了一个路由问题:模型的内部表示包含的正确性信号多于其输出行为所利用的。
英文摘要
Answer correctness is encoded as a recoverable geometric direction in the hidden states of language models. We show that the mean displacement from incorrect to correct answer representations, computed at approximately 70\% of model depth from fifty labeled examples with no parameter updates, yields a scoring direction that outperforms zero-shot log-probability scoring by up to +32.0 percentage points on factual benchmarks (ARC-Challenge and MMLU) and by +38.1 to +51.8 percentage points on TruthfulQA, across five models spanning 1B to 8B parameters in three architecture families (Llama, Qwen, Gemma). The method requires one forward pass and one dot product per candidate; no generation is performed at inference. Applied as a hallucination detector on individual (question, answer) pairs, the recovered direction achieves 0.693~AUROC versus 0.578 for log-probability scoring. We additionally find that correctness directions for factual reasoning, domain knowledge, and calibrated truthfulness are near-orthogonal in representation space, revealing that language models allocate geometrically independent subspaces to qualitatively distinct notions of correct answer, with architecture-dependent variation in the degree of separation. This structure explains the observed transfer pattern---the direction calibrated on factual questions transfers within task type but not across it---and suggests that LLM calibration failures may reflect a routing problem: the model's internal representation contains more correctness signal than its output behaviour exploits.