arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

REFLEX:将MoE推理重新思考为扩散语言模型中的感知细化的计算分配

REFLEX: Rethinking MoE Inference as Refinement-Aware Compute Allocation in Diffusion Language Models

Xiang Xia, Cheng Yan, Yiming Zhang, Jiazheng Liu, Hongyu Zhang, Wuyang Zhang

arXiv 2608.01784首次发表:更新:

AI 中文总结

REFLEX是一种无需训练的MoE推理方法,通过粗到细的专家预算分配和前沿进度分数,在扩散语言模型中平均减少15%专家计算量,同时保持或提升生成质量,实现更优的质量-计算权衡。

AI 中文摘要

混合专家(MoE)模型通过为每个token仅激活一小部分专家来增加参数容量。这种条件计算范式使自回归语言模型能够扩展模型容量,而不会成比例地增加每个token的计算量。然而,在扩散语言模型(DLM)中,尽管不同token的细化需求差异很大,但每个去噪前向过程会共同重新访问所有token位置,而默认的固定token选择路由为它们分配统一的专家预算,导致专家计算与细化需求不匹配。因此,我们认为DLM中的MoE推理应被视为在异构token细化状态下的感知细化计算分配。我们提出REFLEX(REfinement-aware FLEXible expert allocation,感知细化的灵活专家分配),这是一种无需训练的方法,在保持默认路由不变的同时,围绕不断演变的细化过程重组专家计算。具体而言,REFLEX引入了从粗到细的专家预算分配层次结构,使计算与块相对细化角色保持一致,同时使用前沿进度分数解决活跃块的优先级问题。在两个代表性的基于MoE的DLM(LLaDA-MoE和LLaDA2.0-mini)的多个广泛使用的基准测试中,与默认路由相比,REFLEX平均减少了15%的已分配专家计算量,同时在大多数基准测试中保持甚至提高了生成质量。与自回归风格的可变专家路由方法相比,REFLEX还产生了更一致的质量-计算权衡,进一步支持了根据每个去噪前向过程中暴露的异构细化需求分配专家计算的重要性。

英文摘要

Mixture-of-experts (MoE) models increase parameter capacity by activating only a small subset of experts for each token. This conditional-computation paradigm has enabled autoregressive language models to scale model capacity without a proportional increase in per-token computation. In diffusion language models (DLMs), however, each denoising forward jointly revisits all token positions despite their sharply different refinement demands, while the default fixed token-choice routing assigns them a uniform expert budget, creating a mismatch between expert computation and refinement demand. We argue that MoE inference in DLMs should therefore be viewed as refinement-aware compute allocation across heterogeneous token refinement states. We propose REFLEX (\textbf{RE}finement-aware \textbf{FLEX}ible expert allocation), a training-free method that keeps the default router unchanged while reorganizing expert computation around the evolving refinement process. Specifically, REFLEX introduces a coarse-to-fine hierarchy for expert-budget allocation that aligns computation with block-relative refinement roles while using the Frontier-Progress Score to resolve active-block priorities. Across multiple widely used benchmarks on two representative MoE-based DLMs, LLaDA-MoE and LLaDA2.0-mini, REFLEX reduces allocated expert computation by 15\% on average while preserving or even improving generation quality on most benchmarks relative to default routing. Compared with autoregressive-style variable-expert routing methods, REFLEX also yields a more consistent quality--computation trade-off, further supporting the importance of allocating expert computation according to the heterogeneous refinement demands exposed within each denoising forward.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑