arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用语义匹配抑制重新思考特征依赖评估

Revisiting Shape and Texture Reliance with Category-Separability-Calibrated Suppression

Ning Jiang, Tianyi Luo, Zhengyong Huang, Yao Sui

arXiv 2607.16298首次发表:更新:

发表机构

Institute of Medical Technology, Peking University Health Science Center; National Institute of Health Data Science, Peking University; Institute for Artificial Intelligence, Peking University; School of Computer Science and Engineering, Sun Yat-sen University(北京大学医学部医学技术研究所; 北京大学国家卫生数据科学研究院; 北京大学人工智能研究院; 中山大学计算机科学与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究重新审视视觉识别模型特征依赖问题,引入语义匹配评估框架,通过比较不同抑制下性能,发现卷积神经网络更依赖纹理,视觉Transformer更具鲁棒性,指出语义可比性对解释特征依赖很关键,其优势或与人类视觉皮层表征有关。

AI 中文摘要

理解视觉识别模型是否依赖形状、纹理或颜色对于解释其行为至关重要。先前的线索冲突研究强烈影响了人们认为卷积神经网络偏向纹理的观点,但此类测试是在人为冲突下测量线索偏好,而非自然识别过程中的特征依赖。我们通过可控特征抑制重新审视这一问题,发现除非不同抑制操作造成可比的类别级损害,否则性能下降难以解释。我们引入语义匹配评估框架,比较在类别可分性损失匹配水平下的形状和纹理抑制。在此框架下,在ImageNet上训练的卷积神经网络在纹理抑制下比形状抑制下降更明显,揭示出比不匹配抑制分析所显示的更强的纹理依赖性。跨架构比较发现,视觉Transformer在形状和纹理抑制下都比卷积神经网络保持更高准确率。大脑编码进一步表明,在测试抑制设置下,视觉Transformer的表征在神经预测性能上受抑制引起的下降更小。这些发现表明语义可比性对于从抑制实验解释特征依赖至关重要,并表明视觉Transformer的鲁棒性优势可能与更符合人类视觉皮层的表征有关。

英文摘要

Feature-suppression evaluations infer model reliance on shape or texture from the accuracy loss caused by attenuating each type of information. Such losses, however, conflate feature reliance with the amount of category-relevant information removed by the corresponding transformation. Because shape and texture are suppressed using different operators, their effects are not directly comparable. We introduce the Semantic Degradation Index (SDI), which quantifies the suppression-induced reduction in category separability relative to clean images in a fixed clean-reference discriminative space constructed from handcrafted features. On an ImageNet16-like benchmark, we use SDI to compare Gaussian blur for texture suppression with grid distortion for shape suppression over their overlapping degradation range. At comparable SDI values, all five evaluated ImageNet-trained convolutional neural networks (CNNs) retain less accuracy under Gaussian blur than under grid distortion. This results supports stronger texture than shape reliance under the evaluated operators, contrasting with the shape-dominant conclusion obtained from unmatched suppression conditions. The evaluated Vision Transformers (ViTs) also generally retain more accuracy than CNNs under both operators. To determine whether this advantage extends beyond classification, We evaluate fixed brain-encoding models using clean and suppressed images from the Natural Scenes Dataset. Under both operators, ViT features show smaller suppression-induced decreases in noise-ceiling-normalized explained variance than CNN features. These findings establish category separability as an important reference for interpreting suppression-based feature reliance and show that the CNN-ViT robustness difference extends to model representations predictive of human visual cortical responses.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑