arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2608.04084cs.LGcs.CLcs.CV

SpecDrop:无参数类别条件路由的模块化专业化方法

SpecDrop: Parameter-Free Category-Conditioned Routing for Modular Specialization

  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Boyao Wang, Zhihan Lei

AI总结:

本文提出无参数类别条件路由方法SpecDrop,在视觉任务上性能优于参数匹配基线,表明路由的有效性取决于训练信号粒度与类别的对齐而非路由算法。

AI中文摘要:

混合专家(MoE)网络通过学习得到的路由器、门控和负载均衡损失实现专业化,但在总参数预算匹配的情况下,学习得到的路由器性能可能不如等权重无路由基线。瓶颈是路由算法,还是训练信号粒度与目标类别的对齐?本文通过SpecDrop探究该问题,SpecDrop是一种固定的无参数路由方案:K个分支中的每个分支为其分配的类别接收权重pₐ,为其他类别接收小的泄漏权重pᵢ>0,通过与类别无关的固定分母合并,无学习得到的路由参数,也无辅助损失;推理时需要类别标签。在每个图像有一个超类标签的视觉任务(ResNet-110在CIFAR-100上;ViT-S/16在ImageNet-1K上)中,SpecDrop在CIFAR-100上达到79.23%,在ImageNet-1K上达到79.89%,超过不使用标签的参数匹配基线(在CIFAR-100上比密集模型高4.75;在ImageNet-1K上比No-Routing+SE对照高6.53)。这些增益量化了通过路由部署类别监督的效果——并非比基线的标签感知部署更具优势:给定相同标签,仅就准确性而言,掩码密集模型的输出更强(85.2/83.7)。SpecDrop的贡献是将类别标签转化为训练得到的模块化结构:分支-类别对齐率为58%/100%,掩码增益为0.00(CIFAR)/+1.06(ImageNet)——输出空间限制在训练期间被内部化。在训练单元跨多个类别的模糊划分任务(使用30M Transformer的SlimPajama-6B语言建模;使用LoRA在Llama-3.2-1B上的SuperNI指令调优)中,路由机制在种子噪声范围内简化为匹配的无路由对照,这是本文预测的零结果。路由何时有帮助取决于粒度对齐,而非算法选择。代码:this https URL

英文摘要:

Mixture-of-experts (MoE) networks pursue specialization through learned routers, gates, and load-balancing losses, yet at matched total-parameter budgets learned routers can underperform equal-weight No-Routing baselines. Is the bottleneck the routing algorithm, or the alignment between training-signal granularity and the target categories? We probe the question with SpecDrop, a fixed parameter-free routing scheme: each of $K$ branches receives weight $p_a$ for its assigned category and a small leakage $p_i > 0$ otherwise, merged through a category-independent fixed denominator, with no learned routing parameters and no auxiliary losses; the category label is required at inference. On vision tasks where each image has one superclass label (CIFAR-100 on ResNet-110; ImageNet-1K on ViT-S/16), SpecDrop reaches 79.23% on CIFAR-100 and 79.89% on ImageNet-1K, exceeding parameter-matched baselines that do not use the label (+4.75 over dense on CIFAR-100; +6.53 over the No-Routing+SE control on ImageNet-1K). These gains quantify what category supervision buys when deployed through routing -- not an advantage over label-aware deployments of the baselines: given the same label, masking a dense model's outputs is stronger for accuracy alone (85.2 / 83.7). SpecDrop's contribution is converting the label into trained-in modular structure: 68%/100% branch-category alignment, and masking gains of 0.00 (CIFAR) / +1.06 (ImageNet) -- the output-space restriction is largely internalized during training. On fuzzy partitions, where training units span multiple categories (SlimPajama-6B language modeling with a 30M Transformer; SuperNI instruction tuning over Llama-3.2-1B with LoRA), the routing mechanism reduces to the matched No-Routing controls within seed noise, the null our thesis predicts. Granularity alignment, not algorithm choice, localizes when routing helps. Code: https://github.com/Beryex/SpecDrop

补充信息

↑