arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Llama-Mobile:视觉语言模型的高效2.7位量化

Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs

Luka Ribar, Jeevan Bhoot, Douglas Orr

arXiv 2608.21134首次发表:更新:

发表机构

Graphcore Research; Arm(Graphcore研究院; Arm公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLMs部署于移动设备的资源瓶颈,提出结合自生成训练数据的量化流水线与2.7位参数格式的框架,将Llama 3.2 11B Vision Instruct模型压缩至3.7 GB且保留视觉问答任务性能。

AI 中文摘要

将视觉语言模型(VLMs)部署在移动设备上颇具挑战,因为它们需要大量内存和计算资源。我们提出一种用于量化VLMs以在资源受限硬件上高效推理的框架。我们的方法结合了一种量化流水线,该流水线利用模型自身生成训练数据,无需访问原始训练设置,还采用了一种新颖的每参数2.7位格式,支持在Arm CPU上高效执行。我们通过将Llama 3.2 11B Vision Instruct模型压缩至3.7 GB(激活值为8位)来验证该方法,其在一组标准视觉问答任务上保留了较强性能。

英文摘要

Deploying vision-language models (VLMs) on mobile devices is challenging due to their significant memory and compute requirements. We present a framework for quantizing VLMs for efficient inference on resource-constrained hardware. Our approach combines a quantization pipeline that uses the model itself to generate training data and does not require access to the training setup, with a novel 2.7-bit-per-parameter format supporting efficient execution on Arm CPUs. We validate our approach by compressing the Llama 3.2 11B Vision Instruct model to 3.7 GB with 8-bit activations, preserving strong performance on a set of standard visual question answering tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑