AI 中文总结
本研究提出 ASMI 作为大语言模型的不确定性信号,通过掩码注意力头测量子网络间的 BALD 互信息,在基于事实的问答任务中优于基线,可有效识别自信但脆弱的预测。
AI 中文摘要
我们提出,模型对某个 token 的不确定性不仅体现在其输出分布的广度上,还体现在其自信预测在注意力路径受扰动时是否具有脆弱性。我们将此实例化为 ASMI(注意力子网络互信息),这是一种无需训练的估计量,它会掩码注意力头并测量所得子网络间的 BALD 互信息,同时使用语义一致性核来消除表面形式分歧的影响。该信号并非输出置信度的重述:在基于事实的问答任务中,一项折外测试表明,它能提供超出单次通过置信度和熵的错误预测信息,且集中在「自信但脆弱」的预测上,对其进行处理大致可将置信度过滤器保留的错误减半。这种区分具有区域分级性,因此 ASMI 可预测其适用领域:在答案通过提供的上下文传递时表现强劲,而在答案从参数化知识中召回时则受设计限制。Sem-ASMI 可从单个贪心响应中读取信号,无需最强基线所需的随机生成,且在 12 个基于事实的基准-主干设置中有 10 个与语义熵持平或更优。在相同的 12 个设置中,最佳 ASMI 变体(通常是复用基线已抽取的 10 个样本的自适应变体)有 8 个与最强基线持平或领先,其中 3 个在配对检验下显著领先。在参数化问答任务中,所有变体均回落至或低于零成本 MSP 基线,正如预测的那样,且估计值在重复运行中接近确定性。一项头级别分析表明,追踪此边界的并非头级别脆弱性的存在,而是该脆弱性是否与错误耦合。
英文摘要
We propose that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways. We instantiate this as ASMI (Attention-Subnetwork Mutual Information), a training-free estimator that masks attention heads and measures the BALD mutual information among the resulting subnetworks, with a semantic-agreement kernel to discount surface-form disagreement. The signal is not a restatement of output confidence: on grounded QA an out-of-fold test shows it adds error-predictive information beyond single-pass confidence and entropy, concentrated in \emph{confident-but-fragile} predictions, where acting on it roughly halves the retained error of a confidence filter. The distinctness is regime-graded, so ASMI predicts its own domain of applicability, strong where answers are routed through provided context and bounded by design where they are recalled from parametric knowledge. Sem-ASMI reads the signal from a single greedy response, without the stochastic generations the strongest baselines require, and ties or beats Semantic Entropy on ten of the twelve grounded benchmark-backbone settings. Across the same twelve settings, the best ASMI variant, typically the adaptive one reusing the ten samples already drawn for the baselines, ties or leads the strongest baseline in eight, significantly in three under a paired test. On parametric QA all variants revert to or below the zero-cost MSP baseline, exactly as predicted, and the estimates are near-deterministic across reruns. A head-level analysis shows that what tracks this boundary is not the presence of head-level fragility but whether that fragility couples to errors.
Comments19 pages, Under review