arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

缓解推理任务中的幻觉:基于无需训练的、不确定性引导的转向方法

Alleviating Hallucination in Reasoning Tasks with Training-Free Uncertainty-Guided Steering

Litian Liu, Qiqi Hou, Yubing Jian, Reza Pourreza, Mohammad Ghavamzadeh, Roland Memisevic, Yao Qin, Hong Cai

arXiv 2609.38962首次发表:更新:

发表机构

Qualcomm AI Research; UC Santa Barbara(高通人工智能研究院; 加州大学圣塔芭芭拉分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出无需训练的USteer机制,利用置信度梯度调整层激活,在推理时引导生成低不确定性输出,从而减少幻觉并提升准确性。

AI 中文摘要

近期关于大型语言模型幻觉检测的研究表明,对于固定的预训练模型和推理任务,可以估计模型对其输出正确性的置信度。这类不确定性估计主要被用于通过检测或过滤虚构内容来提高真实性。在本工作中,我们探讨这些信号能否被更主动地利用,直接提升模型生成答案的准确性。我们提出USteer,一种简单且无需训练的转向机制,该机制在推理过程中利用置信度度量关于激活值的梯度来调整模型逐层的激活。这一过程在推理时引导生成朝向不确定性更低的输出,而无需修改模型参数或需要额外的监督。我们表明,这种方法在多种任务上持续减少幻觉,证明置信度信号不仅可用于检测,还能有效地对模型行为进行推理时的控制。

英文摘要

Recent work on hallucination detection in large language models has shown that, for a fixed pre-trained model and reasoning task, it is possible to estimate the model's confidence in the correctness of its outputs. Such uncertainty estimates have primarily been used to improve truthfulness by detecting or filtering confabulations. In this work, we ask whether these signals can instead be used more proactively to directly improve the accuracy of model-generated answers. We propose USteer, a simple, training-free steering mechanism that adjusts a model's layer-wise activations during inference using the gradient of a confidence measure with respect to the activations. This procedure nudges generation toward outputs with lower uncertainty at inference time, without modifying model parameters or requiring additional supervision. We show that this approach consistently reduces hallucination across a range of tasks, demonstrating that confidence signals can be leveraged not only for detection, but also for effective inference-time control of model behavior.

CommentsNeurips 2026 main conference paper

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑