arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16656cs.CV

通道级与令牌感知的视觉状态空间对偶后训练量化

Channel-Wise and Token-Aware Post-Training Quantization for Visual State Space Duality

Jonghyeon Lim, Changhoon Yim

首次发表
浏览论文内容

中文总结 AI 辅助

针对视觉状态空间对偶(VSSD)模型,提出通道级令牌平衡输出感知裁剪(CTOAC)量化方法,通过通道级裁剪和令牌平衡损失缓解激活量化瓶颈,在保持ImageNet-1K及下游任务精度的同时实现高达1.42倍加速。

中文摘要 AI 辅助

状态空间模型(SSMs),特别是Mamba,已成为基于注意力架构的高效替代方案,并已通过ViM、VMamba和视觉状态空间对偶(VSSD)扩展到视觉领域。然而,VSSD的低比特后训练量化(PTQ)行为仍未得到充分理解。对VSSD-Tiny进行权重-激活分离分析,识别出激活量化是低比特的主要瓶颈,而选定VSSD骨干线性层的代表性输入表现出强烈的通道级幅度变化和令牌局部极值。我们提出了通道级令牌平衡输出感知裁剪(CTOAC)方法,通过最小化对应线性输出上的令牌平衡重建损失来学习每个输入通道的裁剪边界。仅选定的线性层及其输入激活被量化;其他骨干操作保持原始精度。在VSSD-Tiny、VSSD-Small和VSSD-Base上,所提出的CTOAC方法保持了ImageNet-1K准确率,并在更激进的精度设置下比评估的基线更为稳健。将相同的量化范围应用于COCO和ADE20K上的VSSD骨干,保持了强大的目标检测、实例分割和语义分割性能。优化的RTX 4090部署配置相对于FP32实现了高达1.42倍的端到端加速。

英文摘要

State space models (SSMs), particularly Mamba, have emerged as efficient alternatives to attention-based architectures and have been extended to vision through ViM, VMamba, and Visual State Space Duality (VSSD). Yet the low-bit post-training quantization (PTQ) behavior of VSSD remains insufficiently understood. A weight-activation split on VSSD-Tiny identifies activation quantization as the dominant low-bit bottleneck, while representative inputs to selected VSSD-backbone linear layers exhibit strong channel-wise magnitude variation and token-localized extremes. We propose the Channel-wise Token-balanced Output-Aware Clipping (CTOAC) method, which learns per-input-channel clipping bounds by minimizing a token-balanced reconstruction loss on the corresponding linear outputs. Only the selected linear layers and their input activations are quantized; other backbone operations retain their original precision. Across VSSD-Tiny, VSSD-Small, and VSSD-Base, the proposed CTOAC method retains ImageNet-1K accuracy and remains substantially more robust than the evaluated baselines at more aggressive precision settings. Applying the same quantization scope to VSSD backbones on COCO and ADE20K preserves strong object detection, instance segmentation, and semantic segmentation performance. An optimized RTX 4090 deployment configuration achieves up to 1.42x end-to-end speedup over FP32.

发表机构

  • Konkuk University(建国大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑