轻量级大语言模型剖析
Profiling Lightweight Large Language Models
- Graduate School of Science and Engineering, Saitama University(埼玉大学理工学研究科)
- ITIS, University of Malaga(马拉加大学信息技术与系统研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究轻量级大语言模型在资源受限环境中的情况,提出基于PTME的实验框架,通过硬件级测量联合测量精度、时间、内存和能耗,对一组模型进行测试,发现代理描述符不足,揭示非主导配置,为不同资源下选模型提供指导。
AI中文摘要:
轻量级大语言模型越来越多地在个人电脑上本地部署,并有望在资源受限的边缘和移动环境中发挥越来越重要的作用。在这种情况下,能耗、执行时间和内存使用直接影响实际可用性,但现有的大语言模型效率评估很大程度上依赖于代理描述符,如参数数量或浮点运算次数,且常与任务精度脱节。本文介绍了一个基于PTME的实验框架,用于对轻量级大语言模型推理进行精度感知剖析,通过直接硬件级测量联合测量精度、执行时间、峰值内存使用和能耗。该方法应用于一组有代表性的轻量级大语言模型,在受控桌面平台上的边缘级资源范围内本地执行,使用涵盖代码生成、数学推理和多任务理解的基准测试。研究发现静态代理描述符能很好地近似推理成本,但无法预测精度。收紧资源范围会增加成本但不影响精度,执行时间增加比能耗更强烈,对更大模型惩罚最大。此外,没有单一模型在所有PTME维度上占主导地位,帕累托分析揭示了仅基于准确性或效率评估会隐藏的非主导配置,为在不同资源范围内选择模型提供了实用指导。这些结果表明,仅按大小、浮点运算次数、延迟或准确性选择轻量级大语言模型可能选错部署候选;PTME剖析揭示了以较低物理成本保持有用精度的配置。
英文摘要:
Lightweight large language models (LLMs) are increasingly being deployed locally on personal computers and are expected to play a growing role in resource-constrained edge and mobile environments. In such settings, energy consumption, execution time, and memory usage directly affect practical usability, yet existing evaluations of LLM efficiency largely rely on proxy descriptors such as parameter count or FLOPs, often decoupled from task precision. This paper introduces a PTME-based experimental framework for the precision-aware profiling of lightweight LLM inference, jointly measuring Precision, execution Time, peak Memory usage, and Energy consumption through direct hardware-level measurements. The methodology is applied to a representative set of lightweight LLMs executed locally under edge-class resource envelopes on a controlled desktop platform, using benchmarks spanning code generation, mathematical reasoning, and multi-task understanding. We find that static proxy descriptors approximate inference cost well but fail to predict precision. Tightening the resource envelope increases cost without affecting precision, amplifying execution time more strongly than energy and penalizing larger models the most. Moreover, no single model dominates across all PTME dimensions, and a Pareto analysis reveals non-dominated configurations that would be hidden by accuracy-only or efficiency-only assessments, providing practical guidance for selecting models under different resource envelopes. These results show that selecting lightweight LLMs by size, FLOPs, latency, or accuracy alone can select the wrong deployment candidate; PTME profiling exposes configurations that preserve useful accuracy at lower physical cost.