发表机构
Institute for AI, Data Analysis and Systems (AIDAS); Oxford Brookes University(人工智能、数据分析与系统研究所(AIDAS); 牛津布鲁克斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出XU-RS框架,利用期望梯度归因方法解释随机集语言模型中答案信度宽度的来源,并在MedQA数据集上验证了其有效性,同时探讨了归一化影响与诊断检查。
AI 中文摘要
不确定性估计告诉我们模型有多不确定,但没有说明原因。如果不知道输入的哪些部分影响模型的不确定性,我们就无法判断该不确定性分数是否依赖于与任务相关的输入特征。我们在使用预训练语言模型构建的随机集分类器中研究这一问题。这些分类器将概率分配给单个答案和答案组,为每个答案产生下概率和上概率;这些概率之间的差异称为信度宽度,用于表示因训练数据有限而产生的关于答案的认知不确定性。我们提出XU-RS,一个将答案的信度宽度归因于提供给语言模型的输入标记(单词或词片)的框架。XU-RS使用期望梯度(一种标准的特征归因方法)来估计输入标记如何对信度宽度做出贡献。该框架在MedQA数据集上使用SmolLM3-3B和Llama-2-7B模型进行评估,结果表明,将XU-RS排名较高的标记的嵌入设为零(零掩蔽)比随机选择的标记的零掩蔽引起信度宽度更大的变化。此外,我们展示了归一化可能导致其他答案组影响答案的宽度,揭示了标记归因如何掩盖数值错误,并提供了诊断检查以验证XU-RS排名较高的标记是否有意义地解释了模型的不确定性。
英文摘要
Uncertainty estimates tell us how unsure a model is, but not why. Without knowing which parts of an input influences a model's uncertainty, we cannot tell whether that uncertainty score depends on input features that are relevant for the task. We study this problem in randomset classifiers built using pretrained language models. These classifiers assign probability to individual answers and to groups of answers, producing lower and upper probabilities for each answer; The difference between these probabilities, called credal width, is used to represent epistemic uncertainty about an answer arising from limited training data. We propose XU-RS, a framework that attributes an answer's credal width to the input tokens (words or word pieces) supplied to a language model. XU-RS uses Expected Gradients (a standard feature attribution method) to estimate how input tokens contribute to credal width. The proposed framework is evaluated on a MedQA dataset using SmolLM3-3B and Llama-2-7B models, demonstrating that setting the embedding of a token ranked highly by XU-RS to zero (zero-masking) causes larger changes in credal width than zero-masking randomly selected tokens. In addition, we show that normalisation can cause other answer groups to influence an answer's width, reveal how token attribution can mask numerical errors, and provide diagnostic checks to verify whether a token ranked highly by XU-RS meaningfully explains model uncertainty.
Comments32 pages, 1 figure, 10 tables