缓解捷径学习:纹理惩罚原型网络
Mitigating Shortcut Learning: Texture-Penalized Prototype Networks
浏览论文内容
中文总结 AI 辅助
提出纹理惩罚原型网络(TPPN),通过纹理惩罚分支抑制高频线索,增强形状偏置,将ResNet-50纹理偏差从55.11%降至29.73%,并提升分布外泛化能力。
中文摘要 AI 辅助
标准卷积神经网络(CNN)由于强烈的归纳纹理偏差而表现出严重的性能退化,这种偏差优先考虑局部高频模式而非全局结构形状。这种依赖性导致在纹理变化或环境效应下出现自信的错误分类。为解决这一缺陷,本研究引入了纹理惩罚原型网络(TPPN),一种新颖的架构框架,在不依赖资源密集型增强数据集的情况下转移这种固有偏差。具体而言,纹理惩罚分支(TPB)施加惩罚以抑制局部纹理代理的提取,迫使网络骨干丢弃高频线索并提取纯净的、形状偏置的表示。通过评估来自最终卷积特征的原型超球体内的相似性,该方法强制执行严格的几何约束,将对象视为基本部分的组合以实现稳健分类。在纹理-形状线索冲突数据集和合成噪声基准上的评估证明了这种结构解耦的更强形状偏差。所提出的框架将基线ResNet-50的固有纹理偏差从55.11%降低到29.73%,超越了现成的视觉变换器(ViT-B/16)的纹理抑制能力。此外,该方法在线索冲突条件下表现出稳健的泛化能力,在遇到分布外(OOD)形状时抵抗纹理捷径学习。模型在更高扰动下保持更强的形状准确性。在干净验证数据上,该架构的准确率仅下降0.90个百分点。这为CNN纹理偏差提供了一种结构性的、高效的解决方案。
英文摘要
Standard Convolutional Neural Networks (CNNs) exhibit severe performance degradation due to a strong inductive texture bias that prioritizes local, high-frequency patterns over global structural shapes. This dependency causes confident misclassifications during textural changes or environmental effects. To address this flaw, this study introduces the Texture-Penalized Prototype Network (TPPN), a novel architectural framework that shifts this inherent bias without depending on resource-intensive augmented datasets. Specifically, a Texture-Penalization Branch (TPB) imposes a penalty to suppress the extraction of local texture proxies, forcing the network backbone to discard high-frequency cues and extract purified, shape-biased representations. By evaluating similarities within a prototype-based hypersphere derived from the final convolutional features, the approach enforces strict geometric constraints, treating objects as compositions of essential parts to achieve robust classification. Evaluations on texture-shape cue-conflict datasets and synthetic noise benchmarks demonstrate the stronger shape bias of this structural disentanglement. The proposed framework reduces the inherent texture bias of a baseline ResNet-50 from 55.11% to 29.73%, surpassing the texture-suppression capabilities of an off-the-shelf Vision Transformer (ViT-B/16). Furthermore, the approach demonstrates robust generalization under cue-conflict conditions, resisting textural shortcut learning when encountering Out-of-Distribution (OOD) shapes. The model maintains stronger shape accuracy against elevated perturbations. On clean validation data, the architecture incurs a minimal drop in accuracy of 0.90 percentage points. This provides a structural, efficient solution to CNN texture bias.
发表机构
- Institute for AI Safety and Security(人工智能安全与安全研究院)
- DLR(德国航空航天中心)
机构由 AI 辅助整理,请以论文原文为准。