RPFQ-ViT:用于视觉Transformer极低比特权重的旋转相位帧量化
RPFQ-ViT: Rotated Phase-Frame Quantization for Extremely Low-Bit Weights in Vision Transformers
浏览论文内容
中文总结 AI 辅助
RPFQ-ViT提出旋转相位帧量化方法,在二维相位平面量化成对通道,以极低比特权重保留ViT投影方向,在ImageNet上达到高精度,模型压缩5.4-7.1倍,延迟降低1.4-1.6倍。
中文摘要 AI 辅助
视觉Transformer(ViT)在图像识别和移动视觉应用中取得了强劲性能,但其高维线性投影和注意力计算仍然带来巨大的存储和推理成本。极低比特量化是一种有前景的解决方案,然而ViT经常遭受严重的精度下降,因为传统的实值标量码本与Transformer投影的方向几何特性不匹配。我们提出了RPFQ-ViT,一种旋转相位帧量化方法,在二维相位平面中对成对通道进行量化,使低比特码能更好地保留投影方向,同时通过轻量级缩放恢复幅度。RPFQ-ViT可作为此http URL的即插即用QAT替代方案,且不修改标准的实值注意力、归一化或激活计算图。在ImageNet-1K上,RPFQ-ViT-B/16在W2/A4下达到79.33% Top-1 / 94.48% Top-5准确率,Swin-T在W2/A8下达到79.30% Top-1 / 94.79% Top-5准确率,DeiT-S在W2/A8下达到77.41% Top-1 / 93.11% Top-5准确率。消融实验、相位几何分析和方向保持指标表明,通道配对、可学习旋转、相位锚点学习和残差相位细化均能提高量化质量。我们进一步在原生iOS和Android运行时栈上部署了RPFQ-ViT图像分类模型;使用2比特打包权重,模型大小相对于FP32缩小约5.4-7.1倍,端到端设备端延迟降低1.4-1.6倍。我们代码库中训练的所有ImageNet结果均使用匹配的300轮训练方案,并报告为三次独立运行的平均准确率。这些结果表明,RPFQ-ViT在极低比特ViT的精度、压缩和实际移动部署之间提供了有利的权衡。
英文摘要
Vision Transformers (ViTs) achieve strong performance on image recognition and mobile vision applications, but their high-dimensional linear projections and attention computations still impose substantial storage and inference costs. Extremely low-bit quantization is a promising solution, yet ViTs often suffer severe accuracy degradation because conventional real-valued scalar codebooks are poorly matched to the directional geometry of Transformer projections. We present RPFQ-ViT, a Rotated Phase-Frame Quantization method that quantizes paired channels in two-dimensional phase planes, enabling low-bit codes to better preserve projection directions while recovering magnitude with lightweight scaling. RPFQ-ViT serves as a drop-in QAT replacement for nn.Linear and does not modify the standard real-valued attention, normalization, or activation computation graph. On ImageNet-1K, RPFQ-ViT-B/16 reaches 79.33% Top-1 / 94.48% Top-5 under W2/A4, Swin-T reaches 79.30% Top-1 / 94.79% Top-5 under W2/A8, and DeiT-S reaches 77.41% Top-1 / 93.11% Top-5 under W2/A8. Ablations, phase-geometry analysis, and direction-preservation metrics show that channel pairing, learnable rotation, phase-anchor learning, and residual phase refinement each improve quantization quality. We further deploy RPFQ-ViT image-classification models on native iOS and Android runtime stacks; with 2-bit packed weights, model size shrinks by roughly $5.4$-$7.1\times$ relative to FP32 and end-to-end on-device latency drops by $1.4$-$1.6\times$. All ImageNet results trained in our codebase use a matched 300-epoch recipe and are reported as mean accuracies over three independent runs. These results show that RPFQ-ViT provides a favorable trade-off among accuracy, compression, and practical mobile deployment for extremely low-bit ViTs.
发表机构
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。