arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAGE:混合专家扩散模型中无分类器引导的子空间对齐

SAGE: Subspace Alignment for Classifier-Free Guidance in Mixture-of-Experts Diffusion Models

Boyu Zhang, Yangming Cheng, Ning Zhang, Pengfei Liu, Weijie Li, Yifan Gao, Hangyu Li, Litong Gong

arXiv 2609.34525首次发表:更新:

发表机构

Alibaba Token Hub, Alibaba Group(阿里巴巴集团阿里云通义实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对MoE扩散模型中CFG因分支子空间不一致导致崩溃的问题,提出训练时正则化器SAGE,对齐无条件激活到条件子空间,零推理成本下将漂移降低9.2倍,并在10亿参数模型上提升DPG-Bench性能9.3%。

AI 中文摘要

采用混合专家(MoE)路由的扩散Transformer是扩展生成模型的主要方案。无分类器引导(CFG)对生成质量至关重要,但过高的引导尺度会引发崩溃。我们发现了二者结合时一个此前未被报道的失败模式:两个CFG分支独立路由,因此其实际激活占据不同子空间。无条件写入随后离开条件子空间,而CFG以引导尺度线性放大该残差。我们提出SAGE,一种训练时正则化器,在不限制路由多样性的前提下,将无条件MoE激活对齐到条件子空间,且推理成本为零。玩具实验表明,SAGE将极端漂移大幅抑制了9.2倍。当扩展到10亿参数文生图模型时,SAGE显著提升生成质量,使峰值DPG-Bench性能提升9.3%。大量实验证明SAGE始终优于基线。

英文摘要

Diffusion Transformers with Mixture-of-Experts (MoE) routing are a leading recipe for scaling generative models. Classifier-Free Guidance (CFG) is essential for generation quality, yet excessively high guidance scales trigger collapse. We identify a previously unreported failure mode in their combination: the two CFG branches route independently, so their realized activations occupy different subspaces. The unconditional write then leaves the conditional subspace, and CFG amplifies that residual linearly in the guidance scale. We propose SAGE, a training-time regularizer that aligns unconditional MoE activations to the conditional subspace without restricting routing diversity, at zero inference cost. Toy experiments show that SAGE dramatically suppresses extreme drift by 9.2x. When scaled to a 1B-parameter text-to-image model, SAGE significantly improves generation quality, delivering a 9.3% boost in peak DPG-Bench performance. Extensive experiments demonstrate that SAGE consistently outperforms the baseline.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑