通道专家混合:用于逐点投影的具有输入自适应混合的静态稀疏支持
Mixture of Channel Experts: Static Sparse Supports with Input-Adaptive Mixing for Pointwise Projections
查看机构详情
- School of Computer Science, Ariel University(阿里尔大学计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对MoE应用于卷积网络的结构缺陷,提出MoCE层替代逐点投影,在减少16.7% MAC和延迟的同时,性能匹配或优于基线与现有通道选择方法。
中文摘要 AI 辅助
专家混合(MoE)通过将每个输入路由到一小部分独立参数化的专家来扩展语言模型。我们发现,将这种设计复制到卷积网络中会因一个结构原因而失败:读取相同输入通道的并行卷积专家会学习到几乎相同的滤波器。因此,我们将专家轴从算子复制转移到通道选择。我们引入了通道专家混合(MoCE),这是一种受MoE启发的结构化稀疏通道混合层,用于替代逐点(1×1)通道降维投影。在MoCE中,一个专家是单个输出通道,具有k远小于C个输入通道的学习稀疏支持。所选通道通过softmax组合,其温度针对每个输入进行预测,因此每个专家可以在均值类和最大值类聚合之间切换。一个残差专家总结未选中的通道,并且负载平衡损失保持通道覆盖的完整性。MoCE将成本与C成二次关系的密集投影替换为相对成本按k/C缩放的机制,所预测的节省在测量的挂钟时间中成立。在ImageNet-1K和CIFAR-100上的ResNet骨干、迁移学习、EfficientViT以及强大的现代训练方案中,MoCE与密集基线和先前的通道选择方法相比,性能相当或更优,同时减少了16.7%的MAC和端到端延迟。
英文摘要
Mixture-of-Experts (MoE) scales language models by routing each input through a small set of independently parameterized experts. We show that copying this design into convolutional networks fails for a structural reason: parallel convolutional experts that read the same input channels learn nearly identical filters. We therefore move the expert axis from operator duplication to channel selection. We introduce Mixture of Channel Experts (MoCE), a structured sparse channel-mixing layer, inspired by MoE, that replaces pointwise (1x1) channel-reduction projections. In MoCE, an expert is a single output channel with a learned sparse support of k << C input channels. The selected channels are combined by a softmax whose temperature is predicted per input, so each expert can move between mean-like and max-like aggregation. A residual expert summarizes the unselected channels, and a load-balancing loss keeps channel coverage complete. MoCE replaces a dense projection whose cost is quadratic in C with a mechanism whose relative cost scales as k/C, and the predicted savings hold in measured wall-clock time. Across ResNet backbones on ImageNet-1K and CIFAR-100, transfer learning, EfficientViT, and a strong modern training recipe, MoCE matches or exceeds dense baselines and prior channel-selection methods while reducing MACs by 16.7% and end-to-end latency.