发表机构
Adelaide University(阿德莱德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究审计注意力-残差Transformer中路由熵作为不确定性信号,通过多项检验和敏感性分析发现其无法超越模型置信度,但存在依赖对照的增益,表明路由信息有条件价值。
AI 中文摘要
动态架构在每个预测旁边留下每个样本的路由痕迹,而扩散路由容易被解读为预测不可靠的标志。我们审计了在CIFAR-10/100上从头训练、使用软分箱校准辅助损失的Swin-Tiny和DeiT-Small的注意力-残差(AR)变体中路由熵的这种解读,探究该痕迹是否携带超出模型自身置信度所揭示的正确性信息。三项检查逐步探究此问题:路由信号是否在固定置信度下出现,是否跨训练种子复现,以及留出预测器能否利用它对抗仅输出和打乱痕迹的对照。随后,敏感性审计注入已知大小的效应,并测量探针恢复的效应比例。固定30测试分箱族中的任何测试均未通过多重性校正,名义命中率和边缘结果均未在其兄弟种子中重现。在24个配对运行中,标量路由探针在路由分层校准中未产生合并改进,而熵轮廓探针在二元对数损失和Brier分数上预测正确性优于给定打乱轮廓的相同探针,但不如仅置信度预测器:对打乱痕迹的增益并未转化为对输出的增益。以完整logit向量为条件,相应的比较仍未解决。审计界定了这些未检测结果可被解读的范围:在注入效应为0.010纳特时,轮廓探针恢复了神谕增益的24-59%,而参考保持校正探针在两个CIFAR-100设置中分别恢复了8%和23%,低于我们为将其应用于真实标签而设定的阈值。结果确立了依赖对照的增益和不完整的估计器恢复,而非条件路由信息的缺失。
英文摘要
Dynamic architectures leave a per-example routing trace beside each prediction, and diffuse routing is easy to read as a sign that the prediction is unreliable. We audit that reading for routing entropy in Attention-Residual (AR) variants of Swin-Tiny and DeiT-Small, trained from scratch on CIFAR-10/100 with a soft-binned calibration auxiliary loss, asking whether the trace carries information about correctness beyond what the model's own confidence already reveals. Three checks probe this increment: does a routing signal appear at fixed confidence, does it replicate across training seeds, and can a held-out predictor exploit it against output-only and shuffled-trace controls? A sensitivity audit then injects effects of known size and measures the fraction of each that the probes recover. No test in the fixed 30-test binned family survives multiplicity correction, and neither the nominal hit nor a borderline result recurs in its sibling seeds. Across 24 paired runs a scalar routing probe yields no pooled improvement in routing-stratified calibration, and an entropy-profile probe predicts correctness better than the same probe given shuffled profiles yet worse than a confidence-only predictor in both binary log-loss and Brier score: a gain over shuffled traces does not become a gain over the output. Conditioning on the complete logit vector leaves the corresponding comparison unresolved. The audit bounds how far these non-detections can be read: at an injected effect of 0.010 nats the profile probe recovers 24-59% of the oracle gain, and a reference-preserving correction probe recovers 8% and 23% in the two CIFAR-100 settings, below the threshold we fixed for applying it to real labels. The results establish control-dependent gains and incomplete estimator recovery, not the absence of conditional routing information.
Comments36 pages (9 pages main text)