arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2405.12107cs.CVcs.CL

Imp:面向移动设备的高能力大型多模态模型

Imp: Highly Capable Large Multimodal Models for Mobile Devices

  • Hangzhou Dianzi University(杭州电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhenwei Shao, Zhou Yu, Jun Yu, Xuecheng Ouyang, Lihao Zheng, Zhenbiao Gai, Mingyang Wang, Jiajun Ding

更新

AI总结:

本文系统研究了轻量级多模态模型的架构、训练策略与数据,提出2B-4B规模的Imp模型,其3B版本超越同规模及13B模型,并可在移动端高效推理。

AI中文摘要:

通过利用大型语言模型(LLMs)的能力,近期的大型多模态模型(LMMs)在开放世界多模态理解中展现出了卓越的通用性。然而,它们通常参数量大且计算密集,从而阻碍了其在资源受限场景中的适用性。为此,人们相继提出了几种轻量级LMMs,以最大化受限规模(例如3B)下的能力。尽管这些方法取得了令人鼓舞的结果,但大多数方法仅关注设计空间的一两个方面,且影响模型能力的关键设计选择尚未得到深入研究。在本文中,我们从模型架构、训练策略和训练数据等方面对轻量级LMMs进行了系统研究。基于我们的发现,我们提出了Imp——一个2B-4B规模的高能力LMMs家族。值得注意的是,我们的Imp-3B模型稳定地优于所有现有的同等规模轻量级LMMs,甚至超越了13B规模的最先进LMMs。借助低位量化和分辨率降低技术,我们的Imp模型可部署在Qualcomm Snapdragon 8Gen3移动芯片上,实现约13 tokens/s的高推理速度。

英文摘要:

By harnessing the capabilities of large language models (LLMs), recent large multimodal models (LMMs) have shown remarkable versatility in open-world multimodal understanding. Nevertheless, they are usually parameter-heavy and computation-intensive, thus hindering their applicability in resource-constrained scenarios. To this end, several lightweight LMMs have been proposed successively to maximize the capabilities under constrained scale (e.g., 3B). Despite the encouraging results achieved by these methods, most of them only focus on one or two aspects of the design space, and the key design choices that influence model capability have not yet been thoroughly investigated. In this paper, we conduct a systematic study for lightweight LMMs from the aspects of model architecture, training strategy, and training data. Based on our findings, we obtain Imp -- a family of highly capable LMMs at the 2B-4B scales. Notably, our Imp-3B model steadily outperforms all the existing lightweight LMMs of similar size, and even surpasses the state-of-the-art LMMs at the 13B scale. With low-bit quantization and resolution reduction techniques, our Imp model can be deployed on a Qualcomm Snapdragon 8Gen3 mobile chip with a high inference speed of about 13 tokens/s.

补充信息

↑