BlueLM-V-3B:面向移动设备上多模态大语言模型的算法与系统协同设计
BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices
- vivo AI Lab(vivo人工智能实验室)
- CUHK MMLab(香港中文大学多媒体实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出算法与系统协同设计的BlueLM-V-3B,通过优化动态分辨率与硬件感知部署,在手机上实现2.7B+400M参数规模、24.4 token/s速度及OpenCompass基准66.1分的领先性能。
AI中文摘要:
多模态大语言模型(MLLMs)的出现和日益普及,在改善沟通、促进学习和解决问题等方面,具有增强日常生活各个方面的巨大潜力。手机作为必不可少的日常伴侣,是MLLMs最有效且最易获取的部署平台,能够将其无缝融入日常任务中。然而,由于内存大小和计算能力的限制,在手机上部署MLLMs面临挑战,若不进行广泛优化,难以实现流畅的实时处理。本文提出了BlueLM-V-3B,这是一种专为在移动平台上高效部署MLLMs而量身定制的算法与系统协同设计方法。具体而言,我们重新设计了主流MLLMs采用的动态分辨率方案,并针对硬件感知部署实施了系统优化,以优化手机上的模型推理。BlueLM-V-3B具有以下关键亮点:(1)小规模:BlueLM-V-3B包含具有2.7B参数的语言模型和400M参数的视觉编码器。(2)高速度:在采用4位LLM权重量化的MediaTek Dimensity 9300处理器上,BlueLM-V-3B实现了24.4 token/s的生成速度。(3)强性能:在OpenCompass基准测试中,BlueLM-V-3B在参数量≤4B的模型中获得了66.1的最高平均分,并超越了一系列参数量大得多的模型(例如MiniCPM-V-2.6、InternVL2-8B)。
英文摘要:
The emergence and growing popularity of multimodal large language models (MLLMs) have significant potential to enhance various aspects of daily life, from improving communication to facilitating learning and problem-solving. Mobile phones, as essential daily companions, represent the most effective and accessible deployment platform for MLLMs, enabling seamless integration into everyday tasks. However, deploying MLLMs on mobile phones presents challenges due to limitations in memory size and computational capability, making it difficult to achieve smooth and real-time processing without extensive optimization. In this paper, we present BlueLM-V-3B, an algorithm and system co-design approach specifically tailored for the efficient deployment of MLLMs on mobile platforms. To be specific, we redesign the dynamic resolution scheme adopted by mainstream MLLMs and implement system optimization for hardware-aware deployment to optimize model inference on mobile phones. BlueLM-V-3B boasts the following key highlights: (1) Small Size: BlueLM-V-3B features a language model with 2.7B parameters and a vision encoder with 400M parameters. (2) Fast Speed: BlueLM-V-3B achieves a generation speed of 24.4 token/s on the MediaTek Dimensity 9300 processor with 4-bit LLM weight quantization. (3) Strong Performance: BlueLM-V-3B has attained the highest average score of 66.1 on the OpenCompass benchmark among models with $\leq$ 4B parameters and surpassed a series of models with much larger parameter sizes (e.g., MiniCPM-V-2.6, InternVL2-8B).