CONQuER:基于在线校准代理的硬件感知混合精度量化
CONQuER: Hardware-Aware Mixed-Precision Quantisation with Online-Calibrated Surrogates
浏览论文内容
中文总结 AI 辅助
研究在资源受限硬件上部署深度神经网络的混合精度量化问题,提出CONQuER统一编译器集成基础设施,结合NSGA-II进化算法与双代理预筛选引擎及在线校准器,能发现硬件依赖的帕累托最优配置,大幅提升推理速度和准确率。
中文摘要 AI 辅助
在资源受限硬件上部署深度神经网络依赖混合精度量化(MPQ),但当前部署工具链使该过程严重碎片化。量化通常在前端框架中作为硬件无关的预处理步骤,与生成物理机器代码的下游编译器脱节,导致配置欠佳。此外,通过详尽的硬件在环(HIL)测试评估这些配置因搜索空间过大难以处理。我们提出CONQuER,一种用于硬件感知MPQ的统一编译器集成基础设施。它将量化转移到TOSA级别的编译器管道中,基于编译器支持进行智能配置处理。为在实际编译预算内评估不同模型层的组合搜索空间,CONQuER将NSGA-II进化算法与双代理预筛选引擎结合,该引擎评估理论缓存内存界限和特征空间各向同性以丢弃不可行配置。然后通过IREE仅在硬件上执行最强候选策略,将执行指标输入在线校准器,在校准器在NSGA-II进化搜索期间使代理模型与真实硬件行为对齐。跨移动和笔记本电脑CPU以及服务器GPU的评估表明,最优量化策略依赖于硬件。通过将量化与编译器降低和物理执行相结合,CONQuER发现帕累托最优配置,推理速度比未量化基线快12.19倍,且top-1准确率在未量化基线的1.44%以内。
英文摘要
Deploying deep neural networks on resource-constrained hardware relies on mixed-precision quantisation (MPQ). current deployment toolchains severely fragment this process. Quantisation typically occurs as a hardware-agnostic preprocessing step in front-end frameworks, disconnected from the downstream compilers that generate the physical machine code. This separation leads to suboptimal configurations where assigned bit-widths map poorly to the target machine's heterogeneous hardware execution blocks such as tensor cores and variable-width vector units, incurring severe runtime execution penalties. Furthermore, evaluating these configurations via exhaustive hardware-in-the-loop (HIL) testing is intractable due to the exponentially large search space. We present CONQuER, a unified compiler-integrated infrastructure for hardware-aware MPQ. CONQuER shifts quantisation into the compiler pipeline at the TOSA level, enabling intelligent configuration handling based on compiler support. To evaluate this combinatorial search space of different of model layers within practical compilation budgets, CONQuER couples an NSGA-II evolutionary algorithm with a dual-surrogate prescreening engine. This engine evaluates theoretical cache memory bounds and feature space isotropy to discard non-viable configurations. CONQuER then executes only the strongest candidate policies on hardware via IREE, feeding the execution metrics into an online calibrator. This calibrator aligns the surrogate models with the true hardware behaviour during an NSGA-II evolutionary search. Evaluation across mobile and laptop CPUs, and server GPUs demonstrates that optimal quantisation policies are hardware-dependent. By coupling quantisation with compiler lowering and physical execution, CONQuER discovers Pareto-optimal configurations up to 12.19x faster inference with top-1 accuracy within 1.44% of the unquantised baseline.