发表机构
CIMeC, University of Trento; DISI, University of Trento; Free University of Bozen-Bolzano(特伦托大学认知科学与技术跨学科研究中心; 特伦托大学信息工程与计算机科学系; 波尔扎诺自由大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究以Qwen3-4B为对象,探究大型语言模型的过度自信现象,通过识别转码器特征揭示其默认机制偏向确定性生成,干预不确定性特征可缓解过度自信错误,且相关特征具有跨场景泛化性。
AI 中文摘要
大型语言模型往往表现出过度自信,当证据表明需要权衡或弃权(不执行)时,仍会给出肯定的回答。我们使用操纵逻辑必然性和可能性的受控推理场景,在Qwen3-4B上研究这种行为,涉及三种表达不确定性的方式:言语认知标记、弃权(不执行)和数值置信度分数。我们的结果证实了这种过度自信的倾向,尤其是在提示模型输出数值置信度分数时。在可解释性层面,我们提出了一种方法,可差异化识别负责不确定性和确定性的转码器(transcoder)特征。我们的分析显示,Qwen3-4B的默认机制通过大量共享特征的广泛联盟支持确定性生成,而不确定性则由一小部分专用特征介导的稀疏覆盖来实现。对这些不确定性特征进行干预,既从因果上证明了过度自信背后的这种失衡,也能缓解过度自信错误。同一组特征可在三种不确定性表达设置、语言以及分布外模态任务中泛化。
英文摘要
Large language models tend to overconfidence, giving assertive answers when the evidence suggests hedging or abstention. Using controlled reasoning scenarios that manipulate logical necessity and possibility, we study this behavior in Qwen3-4B, across three ways to express uncertainty: verbal epistemic markers, abstention, and numeric confidence scores. Our results confirm this tendency toward overconfidence, particularly when the model is prompted to output a numeric confidence score. At the interpretability level, we propose a method that differentially identifies transcoder features responsible for uncertainty and certainty. Our analysis reveals Qwen3-4B's default mechanism favors certainty generation through a broad coalition of shared features, while uncertainty is implemented as a sparse override mediated by a small set of dedicated features. Intervening on these uncertainty features both causally proves this imbalance underlying overconfidence and also mitigate overconfident errors. The same set of features generalise across the three uncertainty-expression settings, languages, and an out-of-distribution modality task.