发表机构
Amazon; Northeastern University(亚马逊; 东北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型生成的跨度级不确定性量化问题,提出SPANUQ轻量级探测器,通过混合贝塔分布估计不确定性,经特殊训练方式构建基准。实验显示其在不确定性质量、速度及错误定位上表现出色,且能在多个大语言模型上通用。
AI 中文摘要
不确定性估计对于大语言模型的可靠部署至关重要,也是大语言模型生成中自我优化的基础。现有方法粒度欠佳,token级分数缺乏语义连贯性,序列级分数无法定位错误。本文形式化了跨度级不确定性估计(SLUE)任务,目标是自然的不确定性粒度——语义连贯的文本跨度。为此引入SPANUQ,一种轻量级探测器,将昂贵的多样本推理中的不确定性知识提炼为对大语言模型隐藏状态的单次前向传递。SPANUQ采用DETR风格的跨度解码器,通过混合贝塔分布同时检测跨度并估计其不确定性,用贝塔负对数似然回归和对比排序目标的原则组合进行训练。构建了SPANUQ - BENCH基准。实验表明SPANUQ在跨度级不确定性质量上表现最佳,速度快10 - 20倍,其基于DETR的跨度检测器F1达0.910,超越最佳启发式方法39.4%,能精确错误定位,该框架在两个模型家族的五个大语言模型上具有通用性。
英文摘要
Uncertainty estimation is essential not only for the trustworthy deployment of large language models (LLMs) but also as a foundation for self-refinement in LLM generation. However, existing approaches operate at suboptimal granularities: token-level scores lack semantic coherence, while sequence-level scores fail to localize errors. We formalize Span-Level Uncertainty Estimation (SLUE), a new task that targets the natural granularity for uncertainty: semantically coherent text spans, each conveying a single assessable unit of meaning. To address this task, we introduce SPANUQ, a lightweight (25M parameter) probe that distills the uncertainty knowledge from expensive multi-sample inference into a single forward pass over LLM hidden states. SPANUQ employs a DETR-style span decoder to simultaneously detect spans and estimate their uncertainty via a Mixture of Beta distribution, trained with a principled combination of Beta NLL regression and contrastive ranking objectives. We construct SPANUQ-BENCH, the first span-level uncertainty benchmark comprising 20K prompts, 293K annotated spans, and continuous soft labels derived from multi-sample claim verification. Experiments on five LLM backbones show that SPANUQ consistently achieves the best span-level uncertainty quality , outperforming the strongest probe baseline and all sampling-based methods while being 10~20x faster. Its DETR-based span detector attains 0.910 F1, surpassing the best heuristic by 39.4%, enabling precise error localization that sequence-level methods cannot provide. The same architecture ports to five LLMs spanning two model families, with one probe trained per backbone, and we additionally observe that sequence-level uncertainty is partially decomposable, suggesting that span-level estimation subsumes sequence-level as a special case. The project page is available damon-demon.github.io/SpanUQ.
CommentsAccepted by NeurIPS 2026. The project page is available at https://damon-demon.github.io/SpanUQ.html