发表机构
National Kaohsiung Normal University(国立高雄师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ConCA通过均值与负输入熵配对形成双描述符,结合深度一维卷积MLP生成通道权重,在6个细粒度基准及iNat2021-mini上均优于相关基线,为FGVR轻量型通道注意力提供新方向。
AI 中文摘要
轻量型通道注意力机制广泛应用于图像分类,但在细粒度视觉识别(FGVR)中的有效性仍有限。多数模块通过全局平均池化(GAP)汇总每个通道,该操作捕捉激活幅度却忽略空间浓度,因此空间分布不同但均值相同的通道会得到相同描述符。我们提出浓度感知通道注意力(ConCA),将均值与移位不变负输入熵(NegEnt)配对,该熵通过对取反激活值应用softmax计算,形成联合编码幅度与浓度的双描述符。参数量与通道数呈线性关系的深度一维卷积多层感知机(MLP)将该对映射为逐通道权重。在6个细粒度基准上,ConCA在受控从头训练协议下,优于无注意力、SE-Net、ECA-Net基线及4个更丰富的基于描述符的模块,且在iNat2021-mini上可跨8个骨干网络泛化。这些结果表明,通道描述符及将其映射为注意力权重的逐通道门控,是FGVR中轻量型通道注意力的重要但未被充分探索的方面。
英文摘要
Lightweight channel attention mechanisms are widely used in image classification, yet their effectiveness in fine-grained visual recognition (FGVR) remains limited. Most modules summarize each channel by global average pooling (GAP), which captures activation magnitude but ignores spatial concentration, so channels with different spatial distributions but identical means receive the same descriptor. We propose Concentration-Aware Channel Attention (ConCA), which pairs the mean with a shift-invariant negative-input entropy (NegEnt), computed via a softmax over the negated activations, forming a dual descriptor that jointly encodes magnitude and concentration. A depthwise 1-D convolutional multi-layer perceptron (MLP), whose parameter count is linear in the number of channels, maps the pair to a per-channel weight. On six fine-grained benchmarks, ConCA improves over attention-free, SE-Net, and ECA-Net baselines as well as four richer descriptor-based modules under a controlled from-scratch protocol, and it generalizes across eight backbones on iNat2021-mini. These results indicate that the channel descriptor, together with the per-channel gating that maps it to attention weights, is an important but underexplored aspect of lightweight channel attention in FGVR.