差异驱动门控:U-Net 解码器的自适应特征融合
Difference-Driven Gating: Adaptive Feature Fusion for U-Net Decoder
- Department of Computer Science and Technology, Institute for Artificial Intelligence, BNRist, IDG/McGovern Institute for Brain Research, Tsinghua University(清华大学计算机科学与技术系、人工智能研究院、北京信息科学与技术国家研究中心、IDG/麦戈文脑科学研究院)
- Chinese Institute for Brain Research (CIBR)(中国脑科学研究院)
- School of Computer Science and Technology, Beijing Jiaotong University(北京交通大学计算机与信息技术学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究 U-Net 风格模型中多尺度特征融合问题,提出基于特征差异和熵差异的门控方法,通过从两特征流差异导出注意力权重生成耦合门控映射,实验表明该方法优于现有方法,为特征融合提供新范式。
AI中文摘要:
U-Net 风格模型在众多应用中广泛使用,其关键步骤是用自上而下的解码器重建低层特征,需精确融合高层语义和低层细节。现有基于注意力的融合方法通常单独从自上而下的解码器特征(全局)或其与自下而上编码器特征的相关性(局部)导出注意力权重。本文探索不同范式,从两个特征流的差异中导出注意力权重。提出两种基于差异的门控方法:特征差异门控(FDG),直接用全局和局部特征的绝对差生成自适应门控映射;熵差异门控(EDG),通过信息熵测量各流的表示确定性,用其符号熵差导出注意力权重。两种方法都产生耦合门控映射,同时调制全局和局部特征。在医学图像分割、遥感图像云去除和语音分离等不同任务上的实验表明,两种方法均优于现有基于注意力的融合方法,且 EDG 表现更好。结果为 U-Net 风格结构中的多尺度特征融合提出了新范式。
英文摘要:
The U-Net style models have been widely used in many applications. A critical step in these models is to reconstruct the lower-level features using a top-down decoder. This reconstruction requires precise fusion of high-level semantics and low-level details. Existing attention-based fusion methods typically derive attention weights from the top-down decoder features (global) alone or the correlation between the top-down decoder features and the bottom-up encoder features (local), then modulate the encoder features using these weights. In this work, we explore a different paradigm: deriving attention weights from the difference between the two feature streams. To this end, we propose two difference-based gating approaches: Feature-difference gating (FDG), which directly uses the absolute difference between global and local features to generate adaptive gating maps, and Entropy-difference gating (EDG), which measures the representational certainty of each stream via information entropy and uses their signed entropy difference to derive the attention weights. Both methods produce coupled gating maps that simultaneously modulate the global and local features. Experiments on different tasks including medical image segmentation, remote sensing image cloud removal and speech separation showed that both methods outperformed existing attention-based fusion methods, and EDG performed better. The results suggested a new paradigm for multi-scale feature fusion in the U-Net style structures.