自适应模型压缩(AMC):用于超低功耗变压器推理的显著性驱动资源分配
Adaptive Model Compression (AMC): Saliency-Driven Resource Allocation for Ultra-Low-Power Transformer Inference
浏览论文内容
中文总结 AI 辅助
研究在资源受限边缘设备部署大规模变压器模型的挑战,提出基于令牌重要性动态分配资源的自适应模型压缩框架AMC,实验证明该方法能降低能耗、提高吞吐量,有效延长移动设备电池寿命,精度略有下降。
中文摘要 AI 辅助
在资源受限的边缘设备上部署大规模变压器模型是一项挑战,因为静态推理存在高能量和内存开销,对简单和复杂令牌统一处理。为此提出自适应模型压缩(AMC),这是一个基于令牌重要性动态分配硬件资源的显著性驱动框架。通过多层架构,识别关键高显著性信息进行全精度处理,同时降低次要数据的秩和位宽。实验表明,在45nm CMOS硬件上,AMC使系统能耗降低59.2%,吞吐量提高2.24倍,有效延长移动设备电池寿命,精度仅略有下降3.6%。
英文摘要
Deploying large-scale transformer models on resource-constrained edge devices remains a challenge due to the high energy and memory overhead inherent in static inference, which processes simple and complex tokens with uniform intensity. To address this, we propose Adaptive Model Compression (AMC), a saliency-driven framework that dynamically allocates hardware resources based on token importance. By implementing a multi-tier architecture, our system identifies critical high-saliency information for full-precision processing while aggressively reducing the rank and bit-width of less significant data. Experimental results demonstrate that AMC achieves a 59.2% reduction in system energy and a 2.24x increase in throughput on 45nm CMOS hardware. This approach effectively extends the battery life of mobile devices by utilizing high-definition compute only where necessary, maintaining robust performance with a marginal 3.6% accuracy trade-off.
发表机构
- Apple USA(苹果公司(美国))
机构由 AI 辅助整理,请以论文原文为准。