利用多模态专家混合中的领域专家实现高效适配
Harnessing Domain Specialists in Multimodal Mixture-of-Experts for Efficient Adaptation
- California Institute of Technology(加州理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出ExpertLens方法,利用多模态MoE中涌现的专家语义专门化,通过解码路由器权重识别领域专家并选择性微调,在数学、医学和遥感任务中以更少参数和4倍加速达到或超越全量微调及LoRA性能。
AI中文摘要:
专家混合(MoE)架构通过稀疏计算扩展模型容量,将每个令牌仅路由到一小部分专家。在本工作中,我们探讨这种稀疏性是否会在多模态MoE中引发涌现的内在组织。我们发现,尽管专家并未被显式训练为模块化,它们仍会在跨模态和跨领域间发展出强烈的语义专门化。基于这一结构,我们引入了ExpertLens,一种无需数据的方法,通过将路由器权重解码为语义上有意义的词汇令牌,直接从预训练模型权重中识别领域专家。我们利用这种专门化,通过选择性地微调与目标领域相关的专家,实现高效的多模态适配。在数学、医学和遥感任务中,ExpertLens在仅更新21.7%至47.0%的模型参数并实现平均4.0倍训练加速的情况下,匹配或超越了全量微调的性能,并且在适配性能和训练效率上均优于LoRA。这些结果表明,为效率而引入的稀疏性可以产生语义模块化,这对高效适配直接有用。
英文摘要:
Mixture-of-Experts (MoE) architectures scale model capacity through sparse computation, routing each token through only a small subset of experts. In this work, we explore whether this sparsity gives rise to emergent intrinsic organization in multimodal MoEs. We find that experts develop strong semantic specialization across modalities and domains despite not being explicitly trained for modularity. Building on this structure, we introduce ExpertLens, a data-free method that identifies domain-specialized experts directly from pretrained model weights by decoding router weights into semantically meaningful vocabulary tokens. We leverage this specialization for efficient multimodal adaptation by selectively fine-tuning experts relevant to a target domain. Across math, medical, and remote sensing tasks, ExpertLens matches or surpasses full fine-tuning while updating only 21.7 - 47.0% of model parameters and achieving a 4.0x average training speedup, and outperforms LoRA in both adaptation performance and training efficiency. These results show that sparsity introduced for efficiency can give rise to semantic modularity that is directly useful for efficient adaptation.