arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38956cs.AI

路由探针可以在没有新信息的情况下改进:模型输出之外不确定性的精确零假设审计

Routing Probes Can Improve Without New Information: An Exact-Null Audit of Uncertainty Beyond Model Outputs

Wenhao Liang, Lin Yue, Wei Emma Zhang, Mingyu Guo, Olaf Maennel, Weitong Chen

AI总结:

本研究通过精确零假设审计证明,视觉Transformer的路由信号对正确性探针的改进可能源于检查点选择偏差而非额外信息,并提出了条件置换检验作为更可靠的验证方法。

AI中文摘要:

现代视觉Transformer的路由信号——专家门控、注意力残差权重和停止分数——通常能改进预测模型是否正确性的探针,而这种改进常被解读为路由携带了超出模型输出的错误信息。我们直接检验了这一推断:保持真实的输出-路由配对,我们从在不相交数据上拟合的固定输出-only生成器中重新抽取正确性标签,使得路由在构造上不携带信息。在此精确标签零假设下,宽度匹配的MLP比较仍报告在51.3%的仅置信度评估(308/600)中存在路由增益,而线性比较则报告无增益。在六模型面板上固定每个训练轨迹,并根据验证对数损失而非验证准确率选择检查点,消除了检测(120次中的50次降至0/120,独立实现的探针中120次中的83次降至0/120),表明基于准确率的检查点选择是原因;在所有输出视图下,原始检测率从27.5%(528/1,920)降至零观测检测。修复后的比较不敏感,在两个匹配设置中,每个设置20次重复中检测到约0.005纳特的人工植入信号0/20次,而基于估计路由律的条件置换检验在11/20和10/20中检测到该信号,且在零假设下很少拒绝。在真实正确性标签上,条件分析在五个DeiT注意力残差家族中产生模型相对证据;其中四个在条件律的两种指定变体下持续存在,且没有家族通过额外的噪声准则。拟合更好的探针和测试增量信息是不同的问题,每个都需要各自的验证。

英文摘要:

Routing signals of modern vision transformers -- expert gates, attention-residual weights and halting scores -- often improve probes that predict whether the model is correct, and the improvement is commonly read as evidence that routing carries information about errors beyond the model's outputs. We test this inference directly: keeping real output-routing pairs, we redraw correctness labels from a frozen output-only generator fitted on disjoint data, so that routing is uninformative by construction. Under this exact label null, a width-matched MLP comparison still reports a routing gain in 51.3% of confidence-only evaluations (308/600), while a linear comparison reports none. Holding each training trajectory fixed on a six-model panel and selecting the checkpoint by validation log loss instead of validation accuracy removes the detections (50/120 to 0/120, and 83/120 to 0/120 in an independently implemented probe), identifying accuracy-based checkpoint selection as the cause; across all output views the raw detection rate falls from 27.5% (528/1,920) to zero observed detections. The repaired comparison is not sensitive, detecting an implanted signal of about 0.005 nats in 0/20 replicates in each of two matched settings, whereas a conditional permutation test built on an estimated routing law detects it in 11/20 and 10/20 and rejects rarely under the null. On real correctness labels, the conditional analysis yields model-relative evidence in five DeiT attention-residual families; in four it persists under two specified variants of the conditional law, and no family passes an additional noise criterion. Fitting a better probe and testing for incremental information are different problems, and each needs its own validation.

补充信息

↑