arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03218cs.CV

VisionMX:解锁视觉模型的微缩放训练后量化

VisionMX: Unlocking Microscaling Post-Training Quantization for Vision Models

Elad Dror Cohen, Ofir Gordon, Lior Dikstein, Idan Achituve, Hai Victor Habi

首次发表
浏览论文内容

中文总结 AI 辅助

针对视觉模型微缩放训练后量化误差,提出VisionMX方法,通过优化有界权重舍入和可折叠仿射激活校正,在多项视觉任务上优于直接转换和现有基线。

中文摘要 AI 辅助

微缩放(MX)格式正成为一种支持硬件的训练和推理高效方法。它们将低精度元素与共享块缩放相结合,但其对视觉模型的影响仍未得到充分探索。我们系统地研究了跨视觉模型和任务的训练后MX量化。对直接转换的分析确定了三个误差来源:块缩放表示、某些小型卷积权重张量与非均匀元素网格的对齐不佳,以及非负激活对符号编码的利用不足。这些发现推动了VisionMX的提出,这是一种训练后MX量化方法,可优化有界权重舍入并对激活应用可折叠仿射校正。我们使用多种MX风格格式在图像分类、目标检测、语义分割和低光图像增强中评估了VisionMX。它优于直接转换和所评估的训练后量化基线,在对MX转换最敏感的架构中获得了最大的性能恢复。

英文摘要

Microscaling (MX) formats are emerging as a hardware-supported approach to efficient training and inference. They combine low-precision elements with shared block scales, but their impact on vision models remains underexplored. We systematically investigate post-training MX quantization across vision models and tasks. An analysis of direct conversion identifies three sources of error: block-scale representation, the poor alignment of some small convolutional weight tensors with nonuniform element grids, and the underuse of signed codes by nonnegative activations. These findings motivate VisionMX, a post-training MX quantization method that optimizes bounded weight rounding and applies a foldable affine correction to activations. We evaluate VisionMX across image classification, object detection, semantic segmentation, and low-light image enhancement using several MX-style formats. It improves on direct conversion and the evaluated post-training quantization baselines, with the largest performance recoveries in architectures most sensitive to MX conversion

发表机构

  • Arm AI Research (AAIR), Arm Ltd(Arm AI Research (AAIR),Arm 公司)

机构由 AI 辅助整理,请以论文原文为准。

↑