通过核心扩展路由与统一计算调度加速统一多模态模型
Accelerating Unified Multimodal Models with Core-Expansion Routing and Unified Computation Scheduling
浏览论文内容
中文总结 AI 辅助
本文提出CE-Router方法,通过核心扩展路由与统一计算调度减少统一多模态模型的冗余计算,在保留98.03%密集理解性能的同时实现1.93倍端到端推理加速,提升了质量与效率。
中文摘要 AI 辅助
统一多模态模型可同时支持理解与生成任务,但在token、层及生成时间步上会产生大量冗余计算。通过token重要性探测,研究人员发现了非对称核心扩展结构:理解任务具有稳定的重要性分量,生成任务在很大程度上共享该分量但需要依赖进度的修正。因此,本文提出CE-Router,其使用任务共享的核心评分器和依赖进度的生成扩展,通过生成分解和跨任务核心对齐进行优化。推理时,CE-Router压缩token计算并向统一计算调度(Unified Computation Scheduling)提供学习到的路由信号,该调度协调层跳过、FFN剪枝、扩散头缓存复用及去噪步提前退出。在两种代表性UMM架构上的实验表明,该方法在两类任务上均实现了一致的质量-效率提升,保留了98.03%的密集理解性能,同时实现了1.93倍的端到端推理加速。
英文摘要
Unified multimodal models jointly support understanding and generation, but incur substantial redundant computation across tokens, layers, and generation timesteps. Through token-importance probing, we identify an asymmetric core-expansion structure: understanding exhibits a stable importance component, while generation largely shares this component but requires progress-dependent corrections. We therefore propose CE-Router, which uses a task-shared core scorer and progress-conditioned generation expansions, optimized through generation decomposition and cross-task core alignment. At inference, CE-Router compacts token computation and supplies a learned routing signal to Unified Computation Scheduling, which coordinates layer skipping, FFN pruning, diffusion-head cache reuse, and denoising-step early exit. Experiments on two representative UMM architectures demonstrate consistent quality--efficiency improvements across both tasks, retaining 98.03\% of dense understanding performance with a 1.93$\times$ end-to-end inference speedup.
发表机构
- Xiamen University(厦门大学)
机构由 AI 辅助整理,请以论文原文为准。