发表机构
Elmore Family School of Electrical and Computer Engineering; Purdue University(埃尔莫尔家族电气与计算机工程学院; 普渡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出整体可持续性评分(HSS),对比SLMs与量化LLMs的边缘部署可持续性,发现优化后的量化LLMs综合表现更优,SLMs仍具竞争力。
AI 中文摘要
边缘AI模型的选择通常仅由单一指标驱动,如准确率、延迟、内存、能耗或安全性,然而可部署的语言模型必须平衡这五项指标。本研究旨在回答一个问题:原生训练的小型语言模型(SLMs)与通过后训练量化压缩的大型语言模型(LLMs),哪一个能提供更可持续的边缘部署权衡方案。我们引入了可复现的整体可持续性评分(HSS),该评分围绕三重底线构建:经济支柱对应能力与系统效率,环境支柱对应GPU运行能耗,社会支柱对应有害提示的鲁棒性。我们对5个BF16精度的SLMs和5个LLMs采用不同量化方法(BF16、INT8、NF4 4-bit、GPTQ 4-bit、GGUF Q4),共得到30种实测配置。能力通过5个零样本基准进行评估;效率使用延迟、吞吐量、峰值显存(VRAM)和能耗衡量;安全性通过5个有害提示的攻击成功率近似衡量。Qwen3-30B-A3B/GGUF Q4在组合池中排名第一(93.38),其次是Mistral-Small-24B/GGUF Q4(92.40),而Phi-4-mini/BF16是该池中排名最高的SLM(89.49)。因此,“原生SLM必须是最可持续的边缘选择”这一假设并未得到普遍支持;优化后的量化LLMs可在整体上胜出,而SLMs因资源需求较低仍具竞争力。量化是一种系统层面的选择,而非单调的精度-效率权衡,且HSS相对于其比较池和代理定义具有相对性。
英文摘要
Edge-AI model selection is commonly driven by one isolated metric - accuracy, latency, memory, energy, or safety, even though a deployable language model must balance all five. Our work focuses on answering the question whether na- tively trained small language models (SLMs) or large language models (LLMs) compressed through post-training quantization offer the more sustainable edge- deployment trade-off. We introduce a reproducible Holistic Sustainability Score (HSS) organized around the triple bottom line: an economic pillar for capability and systems efficiency, an environmental pillar for operational GPU energy and a social pillar for harmful-prompt robustness. Five BF16 SLMs and five LLMs under different quantization approaches - BF16, INT8, NF4 4-bit, GPTQ 4-bit, and GGUF Q4 produce 30 measured configurations. Capability is assessed on five zero-shot benchmarks; efficiency uses latency, throughput, peak VRAM and energy; and safety is approximated by attack success rate on five harmful prompts. Qwen3-30B-A3B/GGUF Q4 ranks first in the combined pool (93.38), followed by Mistral-Small-24B/GGUF Q4 (92.40), while Phi-4-mini/BF16 is the highest- ranked SLM in that pool (89.49). Thus, the hypothesis that native SLMs must be the most sustainable edge choice is not supported universally; optimized quantized LLMs can win overall, while SLMs remain competitive through lower resource demand. Quantization is a systems-level choice rather than a monotonic precision- efficiency trade-off and HSS remains relative to its comparison pool and proxy definitions.