MobileAIBench:面向端侧用例的 LLM 与 LMM 基准测试
MobileAIBench: Benchmarking LLMs and LMMs for On-Device Use Cases
- Salesforce AI Research(Salesforce AI研究院)
- Salesforce Mobile Platform(Salesforce移动平台)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出 MobileAIBench 基准测试框架,用于在真实移动设备上系统评估不同规模和量化级别的 LLM 与 LMM 在延迟、资源消耗及信任安全方面的表现,以加速移动端 AI 部署研究。
AI中文摘要:
由于增强隐私、稳定性和个性化方面的优势,在移动设备上部署大型语言模型(LLMs)和大型多模态模型(LMMs)已受到广泛关注。然而,移动设备的硬件限制使得必须使用参数更少的模型以及量化等模型压缩技术。目前,关于量化对各种任务性能(包括 LLM 任务、LMM 任务,以及关键的信任与安全)的影响,理解仍然有限,并且缺乏在移动设备上系统测试这些模型的适当工具。为弥补这些不足,我们提出了 MobileAIBench,一个用于评估面向移动端优化的 LLMs 和 LMMs 的综合基准测试框架。MobileAIBench 在不同规模、量化级别和任务上评估模型,并在真实设备上测量延迟和资源消耗。我们的开源框架分为两部分,包括一个用于在桌面上运行评估的库,以及一个用于端侧延迟和硬件利用率测量的 iOS 应用。我们全面的分析旨在通过提供关于在移动平台上部署 LLMs 和 LMMs 的性能与可行性的见解,加速移动 AI 的研究与部署。
英文摘要:
The deployment of Large Language Models (LLMs) and Large Multimodal Models (LMMs) on mobile devices has gained significant attention due to the benefits of enhanced privacy, stability, and personalization. However, the hardware constraints of mobile devices necessitate the use of models with fewer parameters and model compression techniques like quantization. Currently, there is limited understanding of quantization's impact on various task performances, including LLM tasks, LMM tasks, and, critically, trust and safety. There is a lack of adequate tools for systematically testing these models on mobile devices. To address these gaps, we introduce MobileAIBench, a comprehensive benchmarking framework for evaluating mobile-optimized LLMs and LMMs. MobileAIBench assesses models across different sizes, quantization levels, and tasks, measuring latency and resource consumption on real devices. Our two-part open-source framework includes a library for running evaluations on desktops and an iOS app for on-device latency and hardware utilization measurements. Our thorough analysis aims to accelerate mobile AI research and deployment by providing insights into the performance and feasibility of deploying LLMs and LMMs on mobile platforms.