arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DistMoE:用于分布式指令调优的混合专家模型中无需私有数据排练的路由机制

DistMoE: Private-data Rehearsal-free Routing in Mixture-of-Experts for Distributed Instruction Tuning

Mainak Singha, Niccolò Biondi, Elisa Ricci, Subhankar Roy

arXiv 2608.09907首次发表:更新:

发表机构

University of Trento; Fondazione Bruno Kessler; University of Bergamo(特伦托大学; 布鲁诺·凯塞勒基金会; 贝加莫大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态大语言模型分布式私有数据场景,提出DistMoE混合专家方法,通过公共锚定专家组合阶段实现无需排练的路由,在视觉-语言基准上取得灵活复用与适配的竞争力性能。

AI 中文摘要

多模态大语言模型(Multimodal Large Language Models, MLLMs)已展现出强大的多模态指令跟随能力,但将其适配至多样的视觉-语言领域通常假设存在集中式数据访问和高成本的联合训练,当数据分布在私有、领域特定或权限受限的客户端时,这种设置存在局限性。为此,我们提出DistMoE,一种用于分布式视觉指令调优的混合专家(Mixture-of-Experts, MoE)方法。在语言解码器的每一层,它为公共前馈网络(feedforward network, FFN)增添一个客户端特定的私有FFN专家,以获取领域特定知识。然而,独立的专家训练会导致私有FFN学习到不同规模和量级的表示,使得专家合并变得困难。为减少客户端特定的漂移,我们引入了公共锚定的专家组合阶段,该阶段仅通过各向同性正则化损失,在混合的本地客户端数据和公共数据上更新路由器和轻量级私有投影适配器,从而实现跨客户端无需排练的组合。推理阶段,DistMoE在公共和私有专家上执行模块化路由,无需明确的领域标签即可实现 token 级的领域组合。在多样的视觉-语言基准上的实验表明,DistMoE在保留对客户端特定知识的模块化控制的同时,实现了灵活的专家复用、有效的领域适配和具有竞争力的性能。代码可在该 https URL 获取。

英文摘要

Multimodal Large Language Models (MLLMs) have shown strong multimodal instruction-following ability, but adapting them to diverse visual-language domains typically assumes centralized data access and costly joint training. This is restrictive when data is distributed across private, domain-specific, or permission-limited clients. To this end, we propose DistMoE, a mixture-of-experts (MoE) approach for distributed visual instruction tuning. In each layer of the language decoder it augments the public feedforward network (FFN) with a client-specific private FFN expert, with the goal to acquire domain-specific knowledge. However, independent expert training causes the private FFNs to learn representation of different scale and magnitudes, making merging the experts difficult. To reduce client-specific drift, we introduce a public-anchored expert composition stage that updates only routers and lightweight private projection adapters on a mix of local client data and public data, via an isotropic regularization loss, therefore making it cross-client rehearsal-free composition. During inference, DistMoE performs modular routing over public and private experts, enabling token-wise domain composition without explicit domain labels. Experiments across diverse visual-language benchmarks show that DistMoE enables flexible expert reuse, effective domain adaptation, and competitive performance while preserving modular control over client-specific knowledge. Codes are available at https://github.com/mainaksingha01/DistMoE.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑