发表机构
Cornell Tech; Cornell University(康奈尔科技校区; 康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对VLM边缘部署中微缩放崩溃问题,提出MiX格式反转微缩放范式,以共享尾数分组私有指数,配合MiX-MX框架,在4.5比特下精度不输NVFP4,面积效率提升25%,速度提升2.3-4.5倍。
AI 中文摘要
视觉语言模型(VLM)在边缘设备上的部署受到内存带宽的严重制约,因此需要激进的低于8比特的量化。由于边缘加速器在面积和功耗上受到严格限制,它们需要端到端的量化模型。然而,多模态令牌之间极端的动态范围差距导致标准块格式遭受“微缩放崩溃”,即单个巨大的离群值劫持共享指数,使周围元素下溢并破坏注意力图。为了打破这一瓶颈,我们提出了微逆缩放(MiX),一种新颖的格式,它在数学上反转了微缩放范式:MiX不是将多个尾数分组在一个共享指数下,而是将私有的逐元素指数分组在单个共享尾数下。为了处理不对称的VLM离群值拓扑,我们引入了一种自适应双格式(MiX-MX)推理框架。通过代数方式分解出共享的MiX尾数,该框架映射到定制加速器,用高效的移位器取代乘法器。在多个VLM上进行端到端评估,我们的4.5比特MiX公式在多模态基准上表现出与NVFP4相当或更优的精度。同时,MiX加速器在面积效率上比NVFP4基线提高了25%,与最先进的加速器Focus相比,在模型上实现了2.3-4.5倍的加速和1.4-2.9倍的能耗降低,证明了逆缩放数据通路在高效VLM部署中具有物理上的优越性。
英文摘要
The deployment of Vision-Language Models (VLMs) on edge devices is severely bottlenecked by memory bandwidth, necessitating aggressive sub-8-bit quantization. Since edge accelerators are strictly constrained by area and power, they require end-to-end quantized models. However, the extreme dynamic range gap between multi-modal tokens causes standard block formats to suffer "microscaling collapse," where a single massive outlier hijacks the shared exponent, underflowing surrounding elements and destroying attention maps. To break this bottleneck, we propose Micro-Inverted-Scaling (MiX), a novel format that mathematically inverts the microscaling paradigm: rather than grouping multiple mantissas under one shared exponent, MiX groups private, per-element exponents under a single shared mantissa. To handle asymmetric VLM outlier topologies, we introduce an adaptive dual-format (MiX-MX) inference framework. By algebraically factoring out the shared MiX mantissa, this framework maps to a custom accelerator, replacing multipliers with efficient shifters. Evaluated end-to-end on multiple VLMs, our 4.5-bit MiX formulation exhibits equivalent or superior accuracy on multi-modal benchmarks compared to NVFP4. Simultaneously, the MiX accelerator delivers a 25% improvement in area efficiency over the NVFP4 baseline and a 2.3-4.5x speedup with 1.4-2.9x energy reduction across models compared to the state-of-the-art accelerator Focus, proving the inverted-scaling datapath is physically superior for efficient VLM deployment.
CommentsAccepted to the 59th IEEE/ACM International Symposium on Microarchitecture (MICRO 2026)