发表机构
Adelaide University(阿德莱德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出框引导证据路由(BGER),利用廉价空间监督在不优化归因图的情况下提升事后类激活图的定位性能,实验表明该方法在多个数据集上显著提高MaxBoxAccV2,且收益机制因架构而异。
AI 中文摘要
事后类激活图(CAMs)是检查图像分类器预测背后证据的标准工具,然而普通训练中没有任何机制鼓励这些图在空间上具有适当性。我们研究廉价的空间监督能否在不优化任何归因图的情况下改进分类器自身的预测类Grad-CAM。框引导证据路由(BGER)在框或掩码监督下于最终特征图上训练一个轻量级门控,并通过门控特征路由分类,而Grad-CAM则在门控前的表示上单独计算,因此被评估的图从不进入训练目标。采用二元交叉熵(BCE)路由损失,BGER在CUB-200-2011上将MaxBoxAccV2从0.584提升至0.715,在Stanford Dogs上从0.757提升至0.832,且准确率相当。匹配的对照实验表明,ResNet-50的大部分提升归因于空间监督重塑骨干网络而非路由本身:当分类绕过门控时,大部分改进仍然保留,而通过门控分离梯度使ResNet-50的结果几乎不变。同样的梯度分离在两种DenseNet-121胸部X光设置中保留了大部分增益,但消除了Swin-T上的表观增益,而直接监督CAM达到更强的定位效果但准确率代价更大。总体而言,空间监督可以改进单独评估的事后CAMs,但收益的机制和大小均取决于架构和评估设置。
英文摘要
Post-hoc class activation maps (CAMs) are a standard tool for inspecting the evidence behind an image classifier's predictions, yet nothing in ordinary training encourages these maps to be spatially appropriate. We study whether inexpensive spatial supervision can improve a classifier's own predicted-class Grad-CAM without ever optimizing an attribution map. Box-Guided Evidence Routing (BGER) trains a lightweight gate on the final feature map under box or mask supervision and routes classification through the gated features, while Grad-CAM is computed separately at the pre-gate representation, so the evaluated map never enters the training objective. With a BCE routing loss, BGER raises MaxBoxAccV2 from $0.584$ to $0.715$ on CUB-200-2011 and from $0.757$ to $0.832$ on Stanford Dogs at comparable accuracy. Matched controls attribute most of the ResNet-50 gain to the spatial supervision reshaping the backbone rather than to routing itself: when classification bypasses the gate, most of the improvement remains, and detaching gradients through the gate leaves the ResNet-50 result nearly unchanged. The same detachment preserves most of the gain in two DenseNet-121 chest X-ray settings but removes the apparent gain on Swin-T, and directly supervising the CAM reaches stronger localization at a larger accuracy cost. Overall, spatial supervision can improve separately evaluated post-hoc CAMs, but both the mechanism and the size of the benefit depend on the architecture and the evaluation setting.
Comments24 pages, 12 figures. Appendix included in the main PDF (pages 10-24)