arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

INT8 是否可移植?嵌入式与汽车加速器上量化推理的跨平台测量研究

Is INT8 Portable? A Cross-Platform Measurement Study of Quantized Inference on Embedded and Automotive Accelerators

Yuyeong Shin

arXiv 2609.16085首次发表:更新:

发表机构

Korea Automotive Technology Institute (KATECH)(韩国汽车技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过跨七类硬件的受控测量,揭示INT8量化推理的可移植性在速度、输出一致性和NPU兼容性三方面失败,证明“一次量化,随处部署”在嵌入式与汽车场景中不安全。

AI 中文摘要

八位整数(INT8)训练后量化是边缘部署的默认方案,其背后有一个广泛持有的假设:INT8 以较小且可预测的精度代价换取更快的推理速度,并且一次量化的模型可以移植到任何目标平台。我们通过一项受控测量研究来检验这一假设,该研究覆盖七类硬件——ARM 和 x86 CPU、独立 GPU、NVIDIA Jetson AGX Orin 的 iGPU 及其 NVDLA 核心,以及两个厂商的 NPU(Qualcomm Hexagon HTP、DEEPX DX-M1)——保持 ONNX 工件和量化尺度固定,使得整数内核或指令集架构(ISA)成为唯一的自由变量。可移植性在三个维度上失败。(1)INT8 加速比的正负号由 CPU 的点积指令集(ARM dotprod/SDOT、x86 VNNI)决定:拥有该指令集的核心加速最高达 2.1 倍,而缺乏该指令集的核心在相同模型和运行时下减速 1.7 倍。(2)INT8 输出不可移植,且规律是不变性而非梯度:FP32 预测在每对目标上逐位相同(1000/1000),而 INT8 预测在共享整数内核的两个目标上恰好 1000/1000 一致,在不共享时则为 958-965/1000——无论边界是 CPU 到 CPU 还是 CPU 到加速器,且对保持不变的 top-1 精度不可见。(3)厂商 NPU 主导量化:自带 QDQ 图在一个 NPU 上静默失败(外部尺度被忽略,精度从 0.75 降至 0.005,而编译、分析和运行均无错误),在另一个 NPU 上则大声失败(编译器拒绝该图),因此只有厂商的原生路径才能产生正确的引擎。我们进一步表明,边缘 NPU 的延迟区间由输出/设备到主机传输大小而非计算决定,并通过固定计算扫描定位转换点。我们发布了脚本和 32 份报告。“一次量化,随处部署”对于嵌入式与汽车部署是不安全的,因为在这些场景中,逐输入确定性和冗余至关重要。

英文摘要

Eight-bit integer (INT8) post-training quantization is the default recipe for edge deployment, under a widely held assumption: INT8 makes inference faster at a small, predictable accuracy cost, and a model quantized once can be carried to any target. We test that assumption with a controlled measurement study across seven hardware classes -- ARM and x86 CPUs, a discrete GPU, an NVIDIA Jetson AGX Orin iGPU and its NVDLA cores, and two vendor NPUs (Qualcomm Hexagon HTP, DEEPX DX-M1) -- holding the ONNX artifact and the quantization scales fixed so the integer kernel or ISA is the only free variable. Portability fails on three axes. (1) The sign of the INT8 speedup is set by the CPU's dot-product ISA (ARM dotprod/SDOT, x86 VNNI): cores that have it speed up by up to 2.1x, cores that lack it slow down by 1.7x, for the identical model and runtime. (2) INT8 outputs are not portable, and the rule is an invariance rather than a gradient: FP32 predictions are bit-identical for every pair (1000/1000), while INT8 predictions agree 1000/1000 exactly when two targets share an integer kernel and 958-965/1000 whenever they do not -- independent of whether the boundary is CPU<->CPU or CPU<->accelerator, and invisible to top-1 accuracy, which is preserved. (3) Vendor NPUs own quantization: a bring-your-own QDQ graph fails silently on one NPU (external scales ignored, accuracy 0.75 -> 0.005 while it compiles, profiles and runs without error) and loudly on the other (the compiler refuses the graph), so only the vendor's native path yields a correct engine. We further show that edge-NPU latency regimes are set by output/device-to-host transfer size rather than compute, and locate the transition with a fixed-compute sweep. We release the scripts and 32 reports. "Quantize once, deploy anywhere" is unsafe for embedded and automotive deployment, where per-input determinism and redundancy matter.

Comments22 pages, 3 figures, 8 tables. Artifact:https://github.com/yyshin-katech/embedded-ai-quantization-guide/tree/paper1-v1

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑