AI 中文总结
研究ARM边缘推理中神经网络输出等价类情况,发现INT8 QDQ训练后量化通过H1+H2属性使输出等价,验证了在不同ARM微内核下的位精确一致性,确定x86破坏不变性的机制,指出INT8推理可重现,行为变化轴是精度。
AI 中文摘要
在x86上,内核调度会将同一神经网络的输出分散到硬件上的多个等价类中。我们探讨这种分散情况是否也适用于大多数边缘机器学习实际运行的ARM边缘推理。在四个跨越Cortex-A53、A72和A76的树莓派设备上,在ONNX Runtime CPU下,固定FP32 CNN的输出中无法观察到微架构差异。当硬件固定为Cortex-A76并仅切换执行提供程序时,对于每个CIFAR-10图像,FP32输出都不一致,23个尾数比特的平均剩余精度为14.97。INT8 QDQ训练后量化将两个轴合并为一个等价类。我们将此追溯到QDQ图的一个结构属性H1+H2:离散网格输入使任何卷积调度具有确定性(H1),并且每层边界处的量化线性操作保持该前提条件(H2)。H1+H2预测在运行时可确认地使用不同ARM微内核的情况下,位精确一致性应扩展到生产CNN。我们在TensorFlow Lite与XNNPACK下的MobileNetV2和ResNet50V2上验证了这一点,其中定时证据证实了A76上的SDOT调度和A72上的NEON乘法累加,但每个模型的500个ImageNet图像的每个中间INT32累加器和每个最终输出都是字节相同的。然后我们确定了在x86上破坏相同不变性的特定x86机制,即PMADDUBSW饱和INT16中间值,而ARM没有类似机制。Schlögl等人的发散现象得到了描述而非反驳。对于在异构ARM机群上部署量化CNN的从业者来说,操作结果是直接的。INT8推理是可重现的模式,相关的行为变化轴是精度,而不是微架构。
英文摘要
On x86, kernel dispatch fragments the outputs of the same neural network into many equivalence classes across hardware. We ask whether the same fragmentation governs ARM edge inference, where most edge ML actually runs. Across four Raspberry Pi devices spanning Cortex-A53, A72, and A76 under ONNX Runtime CPU, microarchitecture is not observable in the outputs of a fixed FP32 CNN. Holding hardware constant at Cortex-A76 and switching only the execution provider, FP32 outputs disagree on every CIFAR-10 image with a mean remaining precision of 14.97 of 23 mantissa bits. INT8 QDQ post-training quantization collapses both axes to a single equivalence class. We trace this to a structural property of QDQ graphs that we call H1+H2: discrete-grid inputs make any Conv dispatch-deterministic (H1) and QuantizeLinear at every layer boundary preserves that precondition (H2). H1+H2 predicts that bit-exact agreement should extend to production CNNs under runtimes that confirmably exercise different ARM microkernels. We verify this on MobileNetV2 and ResNet50V2 under TensorFlow Lite with XNNPACK, where timing evidence confirms SDOT dispatch on A76 and NEON multiply-accumulate on A72 yet every intermediate INT32 accumulator and every final output is byte-identical across 500 ImageNet images per model. We then identify the specific x86 mechanism that breaks the same invariant on x86, namely PMADDUBSW saturating INT16 intermediates, which has no ARM analogue. The Schlögl et al. divergence phenomenon is delineated rather than contradicted. For practitioners deploying quantized CNNs across heterogeneous ARM fleets, the operational consequence is direct. INT8 inference is the reproducible mode and the relevant behavioral variation axis is precision, not microarchitecture.
CommentsSubmitted to Journal of Edge Computing; In-review