超越独立优化:多模态边缘智能中的压缩、混合专家路由和量化交互
Beyond Independent Optimization: Compression, MoE Routing, and Quantization Interactions in Multimodal Edge Intelligence
查看机构详情
- Nirma University(尼玛大学)
- Singapore Institute of Technology(新加坡理工学院)
- Marwadi University(马尔瓦迪大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究多模态边缘智能中高效推理受多种因素限制,回顾相关模型进展,指出技术间相互影响不能独立优化,介绍关键设计权衡,引入视频MoE模型诊断方法,强调多方面开放研究方向。
中文摘要 AI 辅助
高效的多模态推理不仅越来越受到模型质量或浮点运算次数的限制,还受到在延迟、内存和能量限制下保存、移动、路由、缓存和量化多模态表示的成本的限制。本文回顾了高效视觉语言和多模态大语言模型的最新进展,涵盖视觉令牌压缩、视频令牌管理、键值缓存优化、混合专家(MoE)路由、低比特量化、边缘部署和硬件感知基准测试。我们认为这些技术不能被视为独立的优化。视觉令牌压缩会改变下游特征分布和MoE路由决策,路由行为会影响专家利用率和量化敏感性,量化的路由器逻辑会影响专家分配,键值缓存策略会决定保留的多模态证据,硬件限制通常会将计算节省转化为内存和通信瓶颈。我们围绕这些交互组织文献,并确定关键的设计权衡,包括准确性与令牌预算、静态与自适应压缩、稀疏路由效率与专家崩溃,以及低比特推理与特定模态退化。最后,我们引入时间路由一致性作为视频MoE模型的诊断方法,并强调在路由感知压缩、跨模态缓存管理、硬件感知协同设计和多模态边缘智能统一基准测试方面的开放研究方向。
英文摘要
Efficient multimodal inference is increasingly constrained not only by model quality or FLOP count, but also by the cost of preserving, moving, routing, caching, and quantizing multimodal representations under latency, memory, and energy constraints. This paper reviews recent advances in efficient vision-language and multimodal large language models, covering visual token compression, video token management, KV-cache optimization, Mixture-of-Experts (MoE) routing, low-bit quantization, edge deployment, and hardware-aware benchmarking. We argue that these techniques cannot be treated as independent optimizations. Visual token compression alters downstream feature distributions and MoE routing decisions, routing behavior affects expert utilization and quantization sensitivity, quantized router logits influence expert assignment, KV-cache policies determine retained multimodal evidence, and hardware constraints often transform computational savings into memory and communication bottlenecks. We organize the literature around these interactions and identify key design trade-offs, including accuracy versus token budget, static versus adaptive compression, sparse routing efficiency versus expert collapse, and low-bit inference versus modality-specific degradation. Finally, we introduce Temporal Routing Consistency as a diagnostic for video MoE models and highlight open research directions in routing-aware compression, cross-modal cache management, hardware-aware co-design, and unified benchmarking for multimodal edge intelligence.