发表机构
Graphcore Research; Arm(Graphcore研究院; Arm公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对VLMs部署于移动设备的资源瓶颈,提出结合自生成训练数据的量化流水线与2.7位参数格式的框架,将Llama 3.2 11B Vision Instruct模型压缩至3.7 GB且保留视觉问答任务性能。
AI 中文摘要
将视觉语言模型(VLMs)部署在移动设备上颇具挑战,因为它们需要大量内存和计算资源。我们提出一种用于量化VLMs以在资源受限硬件上高效推理的框架。我们的方法结合了一种量化流水线,该流水线利用模型自身生成训练数据,无需访问原始训练设置,还采用了一种新颖的每参数2.7位格式,支持在Arm CPU上高效执行。我们通过将Llama 3.2 11B Vision Instruct模型压缩至3.7 GB(激活值为8位)来验证该方法,其在一组标准视觉问答任务上保留了较强性能。
英文摘要
Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constrained hardware. Our approach combines a quantization pipeline that uses the model itself to generate training data and does not require access to the training setup, with a novel 2.7-bit-per-parameter format supporting efficient execution on Arm CPUs. We validate our approach by compressing the Llama 3.2 11B Vision Instruct model to 3.7 GB with 8-bit activations, preserving strong performance on a set of standard visual question answering tasks.