发表机构
University of California, Irvine(加州大学伊文斯分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究低位GEMM在CPU上的适配问题,提出ExaGEMM框架,通过寄存器驻留查找表执行及分析模型共同探索参数化内核和SIMD ISA支持,缩减候选空间,经验证可大幅提升推理延迟,凸显工作负载感知前沿选择对混合精度LLM工作负载的重要性。
AI 中文摘要
低位通用矩阵乘法(GEMM)对于高效的机器学习推理愈发关键,但极低比特执行对传统CPU并不适配。实际部署涵盖多种不同情况,从1/2/4比特权重到不同的激活精度。在固定单指令多数据(SIMD)和寄存器文件预算下,其可行性、复用机会和支持成本各异。我们提出ExaGEMM,这是一个通过寄存器驻留查找表执行实现CPU原生低位GEMM的工作负载感知协同设计和探索框架。关键在于现有SIMD数据路径已涵盖表生成和累加,新硬件只需寄存器内选择/馈送机制并明确建模成本。ExaGEMM利用寄存器可行性、计算成本、内存流量和硬件开销的分析模型共同探索参数化内核和轻量级SIMD指令集架构(ISA)支持,在模拟前将候选空间缩减99.2%。然后识别非支配支持点并生成ISA规范、gem5补丁和GEMM内核进行验证。在代表性机器学习模型和CPU目标上,ExaGEMM比仅软件基线的延迟提高了13.29倍,同时表明工作负载感知前沿选择对混合精度语言模型工作负载尤为重要。
英文摘要
Low-bit GEMM is increasingly central to efficient ML inference, yet very-low-bit execution remains a poor fit for conventional CPUs. Practical deployment spans fragmented regimes-from 1/2/4-bit weights to varying activation precision-whose feasibility, reuse opportunity, and support cost differ under fixed SIMD and register-file budgets, making lightweight CPU support selection a first-class design problem. We present ExaGEMM, a workload-aware codesign and exploration framework for CPU-native low-bit GEMM via register-resident LUT execution. The key insight is that existing SIMD datapaths already cover table generation and accumulation; the only new hardware is an in-register select/feed mechanism with explicitly modeled cost. ExaGEMM co-explores parameterized kernels and lightweight SIMD ISA support using analytical models of register feasibility, compute cost, memory traffic, and hardware overhead, pruning the candidate space by 99.2% before simulation. It then identifies non-dominated support points and generates ISA specs, gem5 patches, and GEMM kernels for validation. Across representative ML models and CPU targets, ExaGEMM improves latency by 13.29x over software-only baselines, while showing that workload-aware frontier selection is especially important for mixed-precision LLM workloads.
CommentsAccepted to ICCAD 2026