arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13095cs.ARcs.AI

MiMo-V2.5系列的全流程推理优化:将混合滑动窗口注意力效率推向极限

Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit

Xiaomi MiMo Team, Anqi Liu, Aoxin Ma, Bo Chen, Bo Yang, Chen Wang, Chen Zhang, Chengda Tang, Chengwei Wang, Chiheng Lou, Depeng Yan, Fuli Luo, Gang Wang, Hailin… 展开作者

Xiaomi MiMo Team, Anqi Liu, Aoxin Ma, Bo Chen, Bo Yang, Chen Wang, Chen Zhang, Chengda Tang, Chengwei Wang, Chiheng Lou, Depeng Yan, Fuli Luo, Gang Wang, Hailin Zhang, Jiale Sun, Kang Zhou, Rui Huang, Shaohui Liu, Shen Huang, Shijie Cao, Shuaishuai Fan, Tianling Zhou, Xiangwei Deng, Xueyang Xie, Xuli Wang, Yingchun Lai, Yu Yang, Yuan Zhang, Zhen Tang, Zhonghua Deng, Zihan Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

针对MiMo-V2.5模型家族,结合Hybrid SWA、MoE和多模态编码器进行全流程推理优化。通过多种策略优化KVCache系统,构建GCache并开发路由器,还优化多模态输入,构成首个有效涵盖该复合架构的大规模LLM生产服务系统。

中文摘要 AI 辅助

我们提出了针对MiMo-V2.5模型家族的全流程推理优化,该模型家族结合了混合滑动窗口注意力(Hybrid SWA)、稀疏专家混合(MoE)和多模态编码器。与全注意力相比,Hybrid SWA理论上能显著减少注意力计算和KVCache存储,但在生产中实现这些收益需要大量工程工作。我们通过分层预取、SWA感知前缀缓存树和专门的放置策略系统地优化KVCache系统,实现严格的$O(W)$ SWA存储和高缓存命中率。我们进一步构建了具有RDMA优化网络的高性能分布式缓存基础设施GCache,并开发了KVCache亲和路由器以减少计算同时保持负载平衡。我们还针对多模态输入进行了优化,包括GPU图像预处理、并行视频解码和多模态缓存共享。这些优化共同构成了首个在生产中有效涵盖Hybrid SWA + MoE + 多模态复合架构的大规模LLM服务系统。

英文摘要

We present a full-pipeline inference optimization for the MiMo-V2.5 model family, which combines Hybrid Sliding Window Attention (Hybrid SWA), sparse Mixture-of-Experts (MoE), and multimodal encoders. While Hybrid SWA can ideally reduce both attention compute and KVCache storage significantly compared to Full Attention, realizing these gains in production requires substantial engineering effort. We systematically optimize the KVCache system with layerwise prefetch, SWA-aware prefix cache trees, and specialized placement strategies, achieving strict $O(W)$ SWA storage and high cache hit rates. We further build GCache, a high-performance distributed cache infrastructure with RDMA-optimized networking, and develop a KVCache-affinity router to reduce computation while preserving load balancing. We also optimize for multimodal inputs, including GPU image preprocessing, parallel video decoding, and multimodal cache sharing. Together, these optimizations constitute the first large-scale LLM serving system in production that efficiently covers the Hybrid SWA + MoE + multimodal composite architecture.

发表机构

  • Xiaomi(小米)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑