基于PyTorch原生栈在服务器CPU上实现小型NLP模型的高效INT8推理
Efficient INT8 Inference of Small NLP Models on Server CPUs with PyTorch Native Stack
- Intel Corporation(英特尔公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究将SmoothQuant集成到TorchAO,优化PyTorch原生栈在服务器CPU上的INT8推理,使小型NLP模型获最高5.8倍吞吐量加速且精度损失可忽略,已上游到PyTorch和TorchAO便于部署。
AI中文摘要:
小型NLP模型,尤其是BERT系列编码器,即便在大语言模型时代,依然在分类、排序、检索等工业工作负载中发挥重要作用。在服务器CPU上,INT8量化能提供极具吸引力的延迟-吞吐量-成本权衡,但用户越来越期望这种加速能直接在原生PyTorch栈中实现。我们将SmoothQuant集成到TorchAO中,并通过TorchInductor的图级融合以及跨oneDNN、AVX512_VNNI、基于AMX的实现的高效INT8 GEMM内核选择,针对英特尔至强CPU优化了生成的推理路径。在BERT、DistilBERT和XLM-RoBERTa基准测试中,该方法相对于FP32基准实现了最高5.8倍的端到端吞吐量加速,且精度损失可忽略不计,部分情况下无可测量的精度损失。我们还通过屋顶线模型进行详细性能分析验证了工作,该实现已上游到PyTorch和TorchAO,可通过原生PyTorch工具实现开箱即用的部署。
英文摘要:
Small NLP models, especially BERT-family encoders, remain important in industrial workloads such as classification, ranking, and retrieval even in the era of large language models. On server CPUs, INT8 quantization offers an attractive latency-throughput-cost trade-off, but users increasingly expect such acceleration to be available directly in the native PyTorch stack. We integrate SmoothQuant into TorchAO and optimize the resulting inference path for Intel Xeon CPUs through graph-level fusion in TorchInductor and efficient INT8 GEMM kernel selection across oneDNN-, AVX512_VNNI-, and AMX-based implementations. Across BERT, DistilBERT, and XLM-RoBERTa benchmarks, the approach delivers up to 5.8x end-to-end throughput speedup with negligible---and in some cases no measurable---accuracy loss relative to the FP32 baseline. We also validated our work by detailed performance analysis with roofline models. The implementation has been upstreamed to PyTorch and TorchAO, enabling out-of-the-box deployment with native PyTorch tooling