arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

U-Space:揭示语言模型中不确定性何时及为何出现

U-Space: Uncovering When and Why Uncertainty Arises in Language Models

Tobias Braun, Nils Loose, Alexander Herzog, Virginia Ceccatelli, Marcus Rohrbach, Thomas Eisenbarth, Lorenzo Cavallaro

arXiv 2610.09087首次发表:更新:

发表机构

Technische Universität Darmstadt; University College London; Universität zu Lübeck; Mila – Quebec Artificial Intelligence Institute; Mohamed bin Zayed University of Artificial Intelligence(达姆施塔特工业大学; 伦敦大学学院; 吕贝克大学; 米拉-魁北克人工智能研究所; 穆罕默德·本·扎耶德人工智能大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出U-Space,一个低维子空间,通过语义锚点构建正交基并投影词元状态,实现无需训练的可解释词元级不确定性量化,在推理基准上优于现有基线。

AI 中文摘要

大型语言模型正以越来越高的风险参与决策。随着其错误后果的日益严重,一个核心问题变得愈发难以忽视:我们能在多大程度上信任单个答案?然而,由于语言模型能够以流畅的解释和权威的语气呈现错误的结论,识别何时应弃权(不执行)仍然困难。不确定性量化旨在通过估计单个预测的可靠性来解决这一脱节问题。然而,许多现有方法需要重复生成或单独训练的组件,且其标量估计无法揭示不确定性在何处产生或如何在推理过程中演变。近期研究还表明,生成长度可能与不确定性估计和正确性密切相关,这引发了一个问题:估计器的预测能力中有多少来自不确定性特定信息,而非仅来自输出长度。机制可解释性提供了一种解决这些局限性的途径,通过将人类可理解的概念与中间模型状态联系起来。基于这一能力,我们引入了U-Space,一个低维子空间,使模型不断演变的不确定性变得可测量且可解释。我们识别出怀疑和确定性的语义锚点,将其去嵌入方向映射回残差空间,并将它们的对比组合成正交基。U-Lens将每个词元状态投影到这些基向量上,生成可解释的词元级不确定性图,可直接检查或聚合成标量不确定性分数。我们的方法不需要正确性标签、重复生成或训练。在推理基准上,其置信度分数在标准和长度控制评估下均优于既有基线,并且比监督估计器更可靠地迁移。代码:此https URL。

英文摘要

Large language models are informing decisions with ever-higher stakes. As the consequences of their errors grow, a central question becomes harder to ignore: how much can we trust an individual answer? Yet recognizing when to defer remains difficult because language models can present incorrect conclusions with fluent explanations and an authoritative tone. Uncertainty quantification seeks to address this disconnect by estimating the reliability of individual predictions. However, many existing methods require repeated generations or separately trained components, and their scalar estimates do not reveal where uncertainty arises or how it evolves during reasoning. Recent work has also shown that generation length can be strongly associated with uncertainty estimates and correctness, raising the question of how much of an estimator's predictive power comes from uncertainty-specific information rather than output length alone. Mechanistic interpretability offers a way to address these limitations by connecting human-interpretable concepts to intermediate model states. Building on this capability, we introduce the U-Space, a low-dimensional subspace that makes a model's evolving uncertainty measurable and interpretable. We identify semantic anchors for doubt and certainty, map their unembedding directions back into the residual space, and combine their contrasts into an orthogonal basis. The U-Lens projects each token state onto these basis vectors, yielding an interpretable token-level uncertainty map that can be inspected directly or aggregated into a scalar uncertainty score. Our approach requires no correctness labels, repeated generations, or training. Across reasoning benchmarks, its confidence score outperforms established baselines under both standard and length-controlled evaluation and transfers more reliably than supervised estimators. Code: https://github.com/s2labres/U-Space.

CommentsCode: https://github.com/s2labres/U-Space

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑