发表机构
Alibaba Token Hub, Alibaba Group; Shanghai Jiao Tong University; Alibaba Cloud Computing(阿里巴巴集团阿里巴巴Token Hub; 上海交通大学; 阿里云计算)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DIET提出基于删除响应的免训练专家剪枝框架,通过最小化总体多样性损失选择保留专家,在LingBot-Video 30B-A3B上剪枝50%专家并提升VBench总分,实现单卡部署。
AI 中文摘要
视频扩散Transformer(DiTs)越来越多地采用混合专家(MoE)架构以减少活跃计算,但其完整的专家存储仍然代价高昂。现有的一次性剪枝标准主要依赖于静态激活或路由统计,无法捕捉专家删除后的层内重路由。我们提出DIET,一种基于删除响应的免训练专家剪枝框架。一次全专家校准过程记录匹配的条件和无条件令牌的专家输出和路由器状态。然后从缓存的张量中重放候选删除,无需额外的模型前向传播。由此产生的删除响应签名通过移除专家时引起的变化来刻画每个专家。DIET通过最小化总体多样性损失(ODL)来选择保留的专家,该损失保留签名空间中的方向覆盖,并将层内局部搜索与层间回归引导的预算搜索相结合,以在各层之间分配专家。在LingBot-Video 30B-A3B上,剪枝50%的专家(从6144个减少到3072个)将检查点从57 GB减少到30 GB,并实现在48 GB GPU上的单卡部署而无需微调。在固定的284例VBench协议下,VBench总分从0.7941提高到0.8115。在测试的各种保留预算下,DIET始终优于从大型语言模型改编的竞争性剪枝基线。
英文摘要
Video diffusion transformers (DiTs) increasingly adopt mixture-of-experts (MoE) architectures to reduce active computation, but their full expert storage remains costly. Existing one-shot pruning criteria mainly rely on static activation or routing statistics and cannot capture layer-level re-routing after expert deletion. We introduce DIET, a training-free expert pruning framework based on deletion responses. A single all-expert calibration pass records expert outputs and router states for matched conditional and unconditional tokens. Candidate deletions are then replayed from cached tensors, requiring no additional model forward passes. The resulting deletion-response signatures characterize each expert by the changes induced when it is removed. DIET selects retained experts by minimizing Overall Diversity Loss (ODL), which preserves directional coverage in signature space, and combines intra-layer local search with an inter-layer regression-guided budget search to allocate experts across layers. On LingBot-Video 30B-A3B, pruning 50% of experts (6,144 to 3,072) reduces the checkpoint from 57 GB to 30 GB and enables single-card deployment on a 48 GB GPU without fine-tuning. Under a fixed 284-case VBench protocol, the VBench Total increases from 0.7941 to 0.8115. Across tested retention budgets, DIET consistently outperforms competitive pruning baselines adapted from large language models.