arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37533cs.CL

E-MoE:用于非因子化扩散语言模型的增强型混合专家

E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models

Arseny Ivanov, Alexander Kolesov, Alexander Korotin, Ivan Oseledets, Mikhail Goncharov

首次发表
浏览论文内容

中文总结 AI 辅助

针对掩码扩散模型少步生成质量受限问题,提出增强型混合专家(E-MoE),利用MoE路由决策作为离散共享潜在变量,构建混合因子化逆向过程,在多个基准上提升少步生成性能。

中文摘要 AI 辅助

掩码扩散模型(MDMs)通过在每个去噪步骤中逐步取消掩码多个标记来生成序列,但其逆向过程通常按位置进行因子化,这限制了在少步数(few-step)机制下的样本质量,而正是在该机制下,扩散相对于自回归解码的速度优势最为重要。最近的一系列工作引入了连续的潜在高斯变量,作为变分自编码器进行训练,以捕捉跨位置的关联,但此类方法容易遭受后验坍塌(posterior collapse),即潜在变量被静默忽略。我们提出了增强型混合专家(E-MoE),它将逆向过程构建为基于混合专家(MoE)骨干网络的专家路由决策所给出的离散共享潜在变量上的因子化分布的混合,且不增加相对于因子化基线的活动参数。在合成多模态基准、二值化MNIST和LM1B上,E-MoE相比因子化基线改善了少步数生成。

英文摘要

Masked diffusion models (MDMs) generate sequences by progressively unmasking several tokens per denoising step, but their reverse process is typically factorized over positions, limiting sample quality in the few-step regime where diffusion's speed advantage over autoregressive decoding matters most. A recent line of work introduces a continuous Gaussian latent, trained as a variational autoencoder, to capture correlations across positions, but such approaches are prone to posterior collapse, where the latent is silently ignored. We propose Enhanced Mixture-of-Experts (E-MoE), which builds the reverse process as a mixture of factorized distributions over a discrete shared latent given by the expert-routing decisions of a Mixture-of-Experts (MoE) backbone, without increasing active parameters over the factorized baseline. Across synthetic multi-modal benchmarks, binarized MNIST, and LM1B, E-MoE improves few-step generation over factorized baselines.

发表机构

  • Applied AI Institute(应用人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑