MiCoPro:面向硬件感知代理模型的端到端混合精度硬件/软件协同设计
MiCoPro: End-to-End Mixed Precision HW/SW Co-design with HW-aware Proxy Model
- Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
MiCoPro是面向边缘AI的混合精度硬件/软件协同设计框架,通过硬件感知代理模型实现延迟约束下的最优量化配置,在两类硬件上达成最高40%延迟降低、精度下降不足3%的效果。
AI中文摘要:
低比特宽度数据的量化神经网络(QNN)已被证明在边缘设备上实现高效存储与计算方面颇具潜力。为缓解精度下降问题同时最大化加速比,逐层混合精度量化(MPQ)成为一种流行解决方案。然而,现有探索MPQ方案的算法在灵活性与效率方面存在局限;传统方法难以理解不同MPQ方案对后训练量化与量化感知训练结果的复杂影响;此外,现有研究中缺少用于MPQ模型优化与部署的端到端框架。为应对这些挑战,我们提出MiCo框架,这是一种面向边缘AI应用的MPQ探索与部署整体框架。该框架采用新型优化算法,在严格延迟约束下搜索精度最优的量化配置;我们进一步将该框架扩展为MiCoPro,其引入鲁棒的硬件感知代理(HAP)模型以提升预测精度与硬件通用性。通过利用目标特定的延迟建模,MiCoPro可实现快速探索,并能直接将PyTorch模型部署为裸机C代码。我们在BitFusion加速器与SIMD扩展的RISC-V处理器上验证了该框架的通用性,实现了最高40%的延迟降低,同时精度下降不足3%。
英文摘要:
Quantized Neural Networks~(QNN) with low-bitwidth data have proven promising in efficient storage and computation on edge devices. To mitigate accuracy degradation while maximizing speedup, layer-wise mixed-precision quantization~(MPQ) becomes a popular solution. However, existing algorithms for exploring MPQ schemes are limited in flexibility and efficiency. Comprehending the complex impacts of different MPQ schemes on post-training quantization and quantization-aware training results is a challenge for conventional methods. Furthermore, an end-to-end framework for the optimization and deployment of MPQ models is missing in existing work. To address these challenges, we propose the MiCo framework, a holistic MPQ exploration and deployment framework for edge AI applications. The framework adopts a novel optimization algorithm to search for accuracy-optimal quantization configurations under strict latency constraints. We further extended the framework to MiCoPro, which introduces a robust Hardware-Aware Proxy (HAP) model to enhance prediction accuracy and hardware versatility. By leveraging target-specific latency modeling, MiCoPro enables rapid exploration and direct deployment from PyTorch models to bare-metal C code. We demonstrate the versatility of our framework on both the BitFusion accelerator and SIMD-extended RISC-V processors, achieving up to 40\% of latency reduction with less than 3\% of accuracy drop.