提示词变体如何影响设备端大语言模型的能耗?
How Do Prompt Variations Affect Energy Consumption in On-Device LLMs?
浏览论文内容
中文总结 AI 辅助
本文探究认知负荷、措辞模式等提示词属性对设备端LLM推理能耗的影响,通过实证研究发现两类属性的不同能耗作用机制,指出需采用感知模型的提示词设计以实现能效优化。
中文摘要 AI 辅助
大语言模型(LLMs)正越来越多地部署在移动设备上,这使得能效成为关键部署约束,但提示词设计对能耗的影响仍未得到充分探索。本文旨在探究认知负荷和措辞模式这两个提示词属性如何塑造设备端LLM推理的能耗行为。我们开展了涵盖提示词属性、数据集、模型和设备的广泛实证研究,采用阶段级分析方法将预填充(prefill)和解码(decode)阶段的能耗分开。我们发现,认知负荷主要影响每token的能耗成本,而措辞模式主要通过token使用量影响能耗。我们的能耗-质量分析进一步表明,提示词设计会在不同模型间以不同方式重塑可实现的性能边界,凸显了在能效设备端LLM推理中需要采用感知模型的提示词设计。代码、数据集和脚本可在该httpsURL获取。
英文摘要
Large language models (LLMs) are increasingly deployed on mobile devices, making energy efficiency a key deployment constraint, yet the energy impact of prompt design remains underexplored. This paper aims to understand how two prompt properties, cognitive load and phrasing pattern, shape the energy behavior of on-device LLM inference. We conduct a broad empirical study covering prompt properties, datasets, models, and devices, with phase-level profiling that separates prefill and decode energy. We find that cognitive load primarily affects the energy cost per token, while phrasing pattern affects energy largely through token usage. Our energy-quality analysis further shows that prompt design reshapes the attainable frontier differently across models, highlighting the need for model-aware prompt design in energy-efficient on-device LLM inference. Code, datasets, and scripts are available at https://amai-gsu.github.io/PromptProperty/.
发表机构
- Georgia State University(佐治亚州立大学)
- Toyota Motor North America(北美丰田汽车公司)
机构由 AI 辅助整理,请以论文原文为准。