AI 中文总结
针对大语言模型在边缘计算环境部署的挑战,提出硬件感知多目标结构化剪枝框架,分两阶段优化,粗粒度移除模块,细粒度搜索最优剪枝率,实验证明能降低模型复杂度,在多方面实现良好权衡,适合边缘部署。
AI 中文摘要
大语言模型因其强大的推理和问答能力而被广泛采用,但在嵌入式和边缘计算环境中部署仍具挑战性。模型剪枝虽能降低规模并保留性能,但联合优化各层、注意力头和多层感知器维度仍很复杂。本文提出硬件感知的多目标结构化剪枝框架,分粗粒度和细粒度两阶段,粗粒度阶段通过多目标深度剪枝移除模块,细粒度阶段用并行贝叶斯优化搜索最优剪枝率。实验表明该方法能在最小影响常识推理任务和零样本性能的情况下降低模型复杂度,在精度、延迟和模型大小间实现良好权衡,适合边缘部署。
英文摘要
Large Language Models (LLMs) have achieved widespread adoption because of their strong reasoning and query-response capabilities. However, deploying them in embedded and edge computing environments remains challenging because of strict latency, memory, and energy constraints. Their large parameter counts and computational demands hinder efficient execution on resource-constrained platforms. Although model pruning has emerged as a viable solution for reducing scale while preserving performance, jointly optimizing layers, attention heads, and Multi-Layer Perceptron (MLP) dimensions remains highly complex. Exhaustively exploring this combined design space is computationally expensive and often leads to local optima or unstable configurations. To address these limitations, we propose a hardware-aware, multi-objective structured pruning framework. The proposed two-stage method explicitly targets latency and model size for efficient deployment on edge devices. In the coarse-grained stage, multi-objective depth pruning removes entire attention and MLP blocks to reduce computational load and memory usage. In the subsequent fine-grained stage, Parallel Bayesian Optimization (PBO) searches for the optimal layer-wise pruning ratios for pruning under latency constraints, while importance-based strategies rank the specific components to be pruned within each layer's allocated budget. Experimental results show that our approach reduces model complexity with minimal impact on commonsense reasoning tasks and zero-shot performance. Our method achieves a favorable trade-off among accuracy, latency, and model size, making it suitable for edge deployment. Across multiple LLMs at 37.5% and 50% pruning ratios, the proposed approach achieves better performance on commonsense reasoning tasks than existing methods while significantly reducing inference cost.
CommentsSubmitted to the ICTAI 2026 (under review)