arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SSDi8:面向状态空间对偶的精准高效8比特量化方法

SSDi8: Accurate and Efficient 8-bit Quantization for State Space Duality

Hyunwoo Kim, Byoungchan Ko, Minseok Kang, Minwoo Kim, Dongjin Lee, Jaehoon Lee, Sungroh Yoon, Dahuin Jung

arXiv 2608.21952首次发表:更新:

发表机构

Chung-Ang University; Soongsil University; Seoul National University(中央大学; 崇实大学; 首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对结构化状态空间对偶(SSD)架构的内存与延迟开销问题,提出首个专为SSD设计的后训练量化框架SSDi8,通过解耦乘法、自适应量化等技术,在保持FP16精度的同时实现最高1.4倍加速,且适配资源受限环境。

AI 中文摘要

序列建模领域的近期进展已凸显出Mamba作为一种状态空间架构,可实现高效的长距离依赖建模,是Transformer的可行替代方案。在此基础上,Mamba-2引入了结构化状态空间对偶(Structured State Space Duality,SSD),其整合了循环与注意力模式,以实现高效性与可扩展性。然而,这种架构扩展大幅增加了内存与延迟开销,凸显出针对SSD定制高效压缩策略的必要性。本研究提出了SSDi8,这是首个专为SSD设计的后训练量化框架,用于维持持久的INT8路径。SSDi8引入了一种重构方案,将逐元素乘法与矩阵乘法解耦,使量化激活值可在各模块间复用。此外,SSDi8在成本效益较高的节点处对通道变化的激活值进行自适应量化,进一步降低延迟。在精度方面,SSDi8明确利用了SSD的内在维度分解,挖掘各轴上不同的异常值分布,并纳入基于逐通道误差统计的误差修正项。综合实验表明,SSDi8实现了与FP16相当的精度,同时在W4A8和W8A8设置下分别实现了最高1.4倍的加速。我们通过在Orin NX设备上部署SSDi8,进一步验证了其在资源受限环境中的鲁棒性。

英文摘要

Recent advances in sequence modeling have highlighted Mamba as a state space architecture offering efficient long-range dependency modeling and providing a viable alternative to Transformers. Building upon this, Mamba-2 introduces the Structured State Space Duality (SSD), which integrates recurrent and attention modes to achieve efficiency and scalability. However, this architectural expansion substantially increases memory and latency overhead, underscoring the need for efficient compression strategies tailored to SSD. In this work, we present SSDi8, the first post-training quantization framework specifically designed for SSD to maintain a persistent INT8 path. SSDi8 introduces a reformulation that decouples element-wise multiplications from matrix multiplications, enabling reuse of quantized activations across modules. Moreover, SSDi8 adaptively quantizes channel-varying activations at cost-effective points, further reducing latency. On the accuracy side, SSDi8 explicitly leverages the intrinsic dimensional decomposition of SSD, exploiting distinct outlier distributions across axes, and incorporates an error correction term based on per-channel error statistics. Comprehensive experiments demonstrate that SSDi8 achieves accuracy comparable to FP16 while delivering up to 1.4x speedup in W4A8 and W8A8 settings. We further validate its robustness in resource-constrained environments by deploying it on the Orin NX device.

CommentsAccepted to ICLR 2026

Journal refInternational Conference on Learning Representations (ICLR), 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑