发表机构
Florida State University(佛罗里达州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大语言模型推理轨迹与答案不一致的问题,提出TrAC框架,结合主动与被动信号实现高效不确定性量化,在数学推理基准上提升了AUROC并降低了AURC。
AI 中文摘要
大语言模型(LLMs)可生成流畅的推理轨迹,却仍可能得出错误答案,因此响应级不确定性估计对弃权(不执行)、人工审核及自适应计算分配至关重要。现有方法大致分为三类:被动单轨迹方法使用token级置信信号;基于采样的方法对比多条完整轨迹,生成成本较高;主动前缀方法探测部分轨迹以研究答案稳定性或偏好转变。然而, none主动从完整推理轨迹中重新引出答案,以衡量其与原始答案的一致性及对原始答案的支持度。为解决这一缺口,我们提出Trace-Conditioned Answer Consistency(TrAC),这是一种正确性监督的不确定性量化框架,结合了锚定在一条完整推理轨迹上的主动与被动信号。其主动组件Prefix-Conditioned Elicitation(PCE)基于完整轨迹重新引出简短答案,并表示其与原始答案的一致性及其token级概率支持;被动组件Trace Uncertainty Profile(TUP)总结原始生成过程中token级不确定性的演变,无需额外解码。随后,一个轻量级头部将这两种表示整合为响应正确性分数。在五个数学推理基准和三个LLM系列上,TrAC相对于八样本自一致性,将宏观AUROC提升1.8%,AURC降低3.4%,仅使用一条完整推理轨迹和一个简短缓存答案探测;当已有八个样本时,用重新引出增强样本共识可进一步将宏观AUROC提升4.3%,AURC降低8.3%,且无需额外完整轨迹生成。
英文摘要
Large language models (LLMs) can generate fluent reasoning traces that nevertheless lead to incorrect answers, making response-level uncertainty estimation important for abstention, human review, and adaptive compute allocation. Existing approaches generally fall into three categories: passive single-trace methods use token-level confidence signals, sampling-based methods compare multiple complete traces at higher generation cost, and active prefix-based methods probe partial traces to study answer stabilization or preference transitions. However, none actively re-elicits an answer from a completed reasoning trace to measure its consistency with and support for the original answer. To address this gap, we introduce Trace-Conditioned Answer Consistency (TrAC), a correctness-supervised uncertainty quantification framework that combines active and passive signals anchored to one completed reasoning trace. Its active component, Prefix-Conditioned Elicitation (PCE), re-elicits a short answer conditioned on the completed trace and represents both its consistency with the original answer and its token-level probabilistic support. Its passive component, Trace Uncertainty Profile (TUP), summarizes how token-level uncertainty evolves throughout the original generation without additional decoding. A lightweight head then integrates the two representations into a response-correctness score. Across five mathematical reasoning benchmarks and three LLM families, TrAC improves macro AUROC by 1.8% and reduces AURC by 3.4% relative to eight-sample self-consistency, while using one complete reasoning trace and a short cached answer probe. When eight samples are already available, augmenting sample consensus with re-elicitation further improves macro AUROC by 4.3% and reduces AURC by 8.3%, without additional full-trace generation.