arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Omni-Diffusion-Distill:统一多模态扩散大语言模型的少步蒸馏

Omni-Diffusion-Distill: Few-Step Distillation of Unified Multimodal Diffusion Large Language Models

Hong Huang, Chenhongyi Yang, Junzhe Sun, Animesh Sinha, Wuyang Chen, Yifan Jiang

arXiv 2610.10990首次发表:更新:

发表机构

Simon Fraser University; Meta(西蒙菲莎大学; Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Omni-Diffusion-Distill是一种两阶段蒸馏框架,可将统一多模态扩散大语言模型的图像生成与多模态理解解码步骤大幅减少,实现高效推理并保持最优性能权衡。

AI 中文摘要

统一多模态扩散大语言模型(dLLM)提供了一种可同时用于图像生成和多模态理解的单一架构,但其迭代解码需要数十到数百次前向传播。现有的少步蒸馏方法大多聚焦于图像生成或文本生成,因此如何将完全离散的多模态dLLM压缩为单一高效学生模型,同时保留生成与理解能力尚不明确。我们提出Omni-Diffusion-Distill,这是一种统一的两阶段蒸馏框架,可在保留强大生成与理解能力的同时,大幅降低统一多模态dLLM的推理成本。Omni-Diffusion-Distill在离散token空间中对齐图像与文本的生成及理解蒸馏。第一阶段,学生模型通过重放缓存的教师轨迹训练以跳过解码步骤;第二阶段,学生模型在自身回滚的中间状态上进行优化。我们进一步通过成对碰撞惩罚(减少并行文本解码下的重复)和熵匹配引导(防止图像生成中拟合锐化教师分布导致的熵崩溃),解决统一蒸馏中的两类退化问题。Omni-Diffusion-Distill在多模态dLLM的解码效率与生成、理解性能间实现了最优权衡,将图像生成解码步骤从128减少至8,多模态理解解码步骤从512减少至64,分别实现18.2倍和21.2倍的 wall-clock 加速。在这些预算下,其文本到图像生成在GenEval上得分为0.828、DPG-Bench上得分为83.0;多模态理解在MM-Vet上的GPT评判得分为20.0,COCO字幕任务上得分为57.2(为教师模型同步骤得分28.4的两倍)。

英文摘要

Unified multimodal diffusion large language models (dLLMs) offer a single architecture for both image generation and multimodal understanding, but their iterative decoding requires tens to hundreds of forward passes. Existing few-step distillation methods largely focus on either image generation or text generation, making it unclear how to compress a fully discrete multimodal dLLM into a single efficient student while preserving both generation and understanding. We introduce Omni-Diffusion-Distill, a unified two-stage distillation framework that retains strong generation and understanding capabilities while substantially reducing the inference cost of a unified multimodal dLLM. Omni-Diffusion-Distill aligns the distillation of both generation and understanding, for both images and text, in the discrete token space. In the first stage, the student is trained to skip decoding steps by replaying cached teacher trajectories, and in the second stage the student is refined on intermediate states along its own rollouts. We further remedy two sources of degradation in unified distillation with a pairwise collision penalty that reduces repetition under parallel text decoding, and entropy-matched guidance that prevents entropy collapse caused by fitting the sharpened teacher distribution in image generation. Omni-Diffusion-Distill achieves state-of-the-art trade-offs between decoding efficiency and generation and understanding performance for multimodal dLLMs, reducing image generation from 128 to 8 decoding steps and multimodal understanding from 512 to 64, giving 18.2x and 21.2x wall-clock speedups. Under these budgets, it scores 0.828 on GenEval and 83.0 on DPG-Bench for text-to-image generation, while reaching GPT judge scores of 20.0 on MM-Vet and 57.2 on COCO captioning (twice the teacher's 28.4 at the same steps) for multimodal understanding.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑