发表机构
University of California, Irvine(加利福尼亚大学欧文分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对CPU上大语言模型推理的量化问题,提出PolyQ方法,通过激活感知的逐通道比特分配及编译时模型编译器优化,实现分数比特部署,在多个模型上提升质量,在不同CPU上有良好性能,证明其在边缘设备推理的实用性、可预测性和节能性。
AI 中文摘要
CPU是设备上大语言模型推理最通用的目标,但现有低比特量化方法要么提供粗糙的操作点,要么提供难以在CPU上高效执行的细粒度混合精度。我们提出了PolyQ,这是一种面向CPU的编译器/量化协同设计,用于在用户指定的平均比特预算下进行激活感知的逐通道比特分配。PolyQ从{2,3,4,8,16}中分配每通道比特宽度,然后使用编译时模型编译器对通道进行排列和聚类成比特均匀的块,生成与SIMD和查找表兼容的内核,并跨运算符合并兼容排列以避免运行时路径上的布局正则化。这将细粒度预算拟合转变为仅用于CPU推理的实用分数比特部署方法。在WikiText-2上的Falcon-H1-3B、Llama2-13B和Qwen3-32B上,PolyQ在3-6比特提供稳定的质量缩放,并在3比特目标下比先前方法将困惑度提高2.4-32.1%。在三个代表性CPU(工作站、笔记本电脑和移动设备)上的端到端测量表明,编译器布局正则化将激活重排序流量减少高达70.8%,预填充延迟和解码吞吐量几乎与配置的比特预算成比例缩放,并且相对于优化的基于查找表的后端,能量/令牌开销保持在2%以下。这些结果表明,分数比特CPU部署在不同边缘目标上是实用的、可预测的且节能的。
英文摘要
CPUs are the most universal target for on-device LLM inference, but existing low-bit quantization methods offer either coarse operating points or fine-grained mixed precision that is difficult to execute efficiently on CPUs. We present PolyQ, a CPU-oriented compiler/quantization co-design for activation-aware channel-wise bit allocation under a user-specified average-bit budget. PolyQ assigns per-channel bit-widths from $\{2,3,4,8,16\}$, then uses a compile-time model compiler to permute and cluster channels into bit-homogeneous blocks, generate SIMD- and LUT-compatible kernels, and merge compatible permutations across operators to keep layout regularization off the runtime path. This turns fine-grained budget fitting into a practical fractional-bit deployment method for CPU-only inference. Across Falcon-H1-3B, Llama2-13B, and Qwen3-32B on WikiText-2, PolyQ provides stable quality scaling from 3--6\,b and improves perplexity by 2.4--32.1\% over prior methods at a 3\,b target. End-to-end measurements on three representative CPUs -- workstation, laptop, and mobile -- show that compiler layout regularization reduces activation reorder traffic by up to 70.8\%, prefill latency and decode throughput scale nearly proportionally with the configured bit budget, and energy/token overhead stays below 2\% relative to an optimized LUT-based back-end. These results show that fractional-bit CPU deployment is practical, predictable, and energy-efficient across diverse edge targets.
CommentsAccepted to ICCAD 2026