AI 中文总结
本工作提出MobiBench,一个统一基准测试套件,用于在多种硬件上评估边缘优化LLMs的性能,涵盖用户和系统指标,并揭示模型与硬件间的权衡,为设备上LLM评估奠定基础。
AI 中文摘要
随着在移动和边缘设备上部署大型语言模型(LLMs)的需求日益增长,理解LLMs在现实世界资源约束下的性能已变得更为关键。尽管存在多种用于设备上推理的执行框架,但本工作的唯一焦点是评估该http URL作为运行时。本工作提出了一个统一的基准测试套件,使用该http URL在各种硬件设置下全面评估边缘优化LLMs的性能。该基准包括面向用户的指标(任务特定准确性、预填充速度、解码速度、首个令牌时间等)和系统级测量(内存消耗、电池消耗等),并涵盖了一系列相当多样化的代表性自然语言任务,如摘要生成和问答。本文通过系统地评估商业系统级芯片(SoCs)和CPU/GPU平台上的几个轻量级模型,提出了一项比较分析,识别出由模型架构和硬件特性导致的重要权衡。为了指导模型优化、运行时开发和边缘设备硬件设计在有效的大规模语言智能方面的未来发展,该基准为使用单一广泛使用的运行时对设备上LLMs进行可重复和全面的评估奠定了基础。
英文摘要
Understanding how large language models (LLMs) perform under real-world resource constraints has become more crucial due to the growing demand for deploying LLMs on mobile and edge devices. While there are a number of execution frameworks for on-device inference, the evaluation of llama.cpp as the runtime is the sole focus of this work. This work, presents a unified benchmarking suite that uses llama.cpp to thoroughly evaluate the performance of edge-optimized LLMs in a variety of hardware settings. The benchmark includes both user-facing metrics (task-specific accuracy, prefill speed, decode speed, time-to-first-token, etc.) and system-level measurements (memory consumption, battery consumption, etc.) and covers a fairly varied set of representative natural language tasks, such as summarization and question-answering. This paper present a comparative analysis that identifies important trade-offs resulting from model architecture and hardware features by methodically assessing several lightweight models on commercial system-onchips (SoCs) and CPU/GPU platforms. In order to guide future developments in model optimization, runtime development, and edge-device hardware design for effective large-scale language intelligence, the benchmark lays the groundwork for repeatable and thorough evaluation of on-device LLMs using a single, widely used runtime.