AI 中文总结
该研究利用Jetson设备的NVJPEG硬件单元,结合多实例方法并行化CPU、GPU等资源,优化视觉模型数据预处理,在大图像尺寸下较最优设计获最高30.02%加速,为边缘推理流程优化提供指南。
AI 中文摘要
数据预处理是边缘设备深度学习工作流的关键组成部分。然而,对JPEG格式存储的数据进行解码的计算量极大,占据了预处理流程的主要部分。因此,提高解码速度对于提升整体吞吐量至关重要,尤其是对于输入图像尺寸较大的情况,这类场景常面临预处理瓶颈。另一方面,边缘设备配备了专用硬件单元以加速媒体处理和图像解码,例如NVIDIA Jetson平台拥有专用的NVJPEG单元,这些单元可用于提升预处理流程的性能。本文介绍了利用此类特定硬件加速单元来分担解码任务的方法,结合多实例方法,实现了对Jetson设备中所有计算资源(包括CPU、NVJPEG、GPU和DLA)的并行化。在本研究中,我们对比了多种潜在的流程设计,针对ResNet18、ResNet50和ResNet152这三种不同规模的模型,评估了批量大小、图像尺寸的影响,以及GPU/DLA推理的特性。最后,我们针对多实例设计开展了微调实验。与未使用该硬件单元的最优化设计相比,引入专用硬件解码单元的多实例设计在大图像尺寸下可实现高达30.02%的加速。基于这些发现,我们证明了在深度学习工作流中使用NVJPEG单元的益处,并提供了用于调整和优化边缘推理工作流的指南。
英文摘要
Data preprocessing is a crucial part of deep learning workflows on edge devices. However, decoding data saved in JPEG format is very compute-intensive and occupies a major portion of the preprocessing pipeline. Therefore, increasing the decoding speed is vital for improving overall throughput, especially for inputs with large image sizes, which are often subject to preprocessing bottlenecks. On the other hand, edge devices are equipped with specialized hardware units to accelerate media processing and image decoding. For instance, the NVIDIA Jetson platform possesses a dedicated NVJPEG unit. These units can be used to enhance the performance of the preprocessing pipeline. This paper introduces the utilization of such specific hardware acceleration units for offloading decoding tasks. By combining this with a multi-instance approach, it allows for the parallelization of all compute resources including CPU, NVJPEG, GPU, and DLA in Jetson devices. In this work, we compare various potential pipeline designs. On ResNet18, ResNet50, and ResNet152, three models with different sizes, we evaluate the impact of batch sizes and image sizes, as well as the characteristics of GPU/DLA inference. Finally, a fine-tuning experiment for multi-instance design has been conducted. The multi-instance design with a specific hardware decoding unit involved offers up to 30.02% speedup for large image sizes, compared with the most optimized design without it. Based on these findings, we demonstrate the benefits of using the NVJPEG unit in deep learning workflows and provide guidelines for tuning and optimizing edge inference workflows.