发表机构
Cognizant AI & Analytics; Cognizant AI Lab; Southern Methodist University (SMU); University of Maryland, Baltimore County(高知特人工智能与分析部门; 高知特人工智能实验室; 南卫理公会大学; 马里兰大学巴尔的摩县分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RotaryQuant通过三轴压缩系统及IsoQuant等技术,在消费级硬件上实现1200亿参数MoE模型的高效运行,内存占用低且性能损失极小。
AI 中文摘要
拥有260亿至1200亿参数的大型混合专家(MoE)语言模型,受三个同时存在的内存压力影响,超出消费级设备的内存容量:驻留权重矩阵、随上下文线性增长的键值(KV)缓存状态,以及必须按需分页的数十个专家子层。本文提出RotaryQuant,这是一个三轴压缩系统,可解决上述三个问题。混合精度权重量化根据架构角色分配位宽:密集层采用4位,路由专家采用2位,共享专家因激活峰度高难以进行激进压缩,故采用8位。LRU专家卸载机制在真实内存压力下,将非驻留专家分页至磁盘。核心创新在于IsoQuant,这是一种KV缓存压缩方法,先应用沃尔什-哈达玛变换,再应用块对角SO(4)旋转,使激活分布各向同性后进行3位标量量化,其运算量为O(d log d),每头需存储256个参数,而密集旋转方法则为O(d²)及16384个参数。融合四内核Metal GPU流水线直接对打包的3位张量执行注意力计算,无需实例化全精度KV状态——这是一种不同的执行模型,而非仅量化方案。该组合系统可在16GB内存预算内适配Gemma 4-26B-A4B和Qwen3-30B-A3B,在32GB内存内适配Nemotron-H 120B,以9至19 token/秒的速度交互式运行,困惑度下降近乎为零(ΔPPL≤+0.0012),在32K上下文长度下检索准确率达100%。
英文摘要
Large mixture-of-experts (MoE) language models with 26--120 billion parameters exceed the memory capacity of consumer devices through three simultaneous pressures: resident weight matrices, key-value (KV) cache state that grows linearly with context, and dozens of expert sublayers that must be paged on demand. We present RotaryQuant, a three-axis compression system that addresses all three. Mixed-precision weight quantization assigns bit-widths by architectural role: 4-bit for dense layers, 2-bit for routed experts, and 8-bit for the shared expert whose high activation kurtosis resists aggressive compression. LRU expert offloading pages non-resident experts to disk under genuine memory pressure. The novel axis is IsoQuant, a KV cache compression method that applies a Walsh--Hadamard transform followed by block-diagonal SO(4) rotations to isotropize activation distributions before 3-bit scalar quantization, requiring $O(d \log d)$ operations and 256 stored parameters per head versus $O(d^2)$ and 16{,}384 for dense rotation methods. A fused four-kernel Metal GPU pipeline performs attention directly on packed 3-bit tensors without materializing full-precision KV state---a different execution model, not just a quantization scheme. The combined system fits Gemma 4-26B-A4B and Qwen3-30B-A3B within a 16\,GB budget and Nemotron-H 120B within 32\,GB, running interactively at 9--19 tok/s with near-zero perplexity degradation ($Δ$PPL $\leq +0.0012$) and 100\% retrieval accuracy at 32K context.
Comments9 Pages, 9 Tables, Initial submission into NeurIPS 2026