发表机构
Jetson-aware embedded deep learning inference(Jetson-aware嵌入式深度学习推理)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对自动驾驶车辆在异构边缘GPU上部署视觉Transformer的问题,提出硬件感知的H-FraDS方法,通过固定调度比率路由帧,适配组件,结合OFA,实现高效推理,提升了帧率、吞吐量并满足实时操作要求。
AI 中文摘要
物理AI系统,如自动驾驶车辆和智能机器,需要基于Transformer的感知模型来满足严格的边缘延迟和能量限制。然而,异构边缘GPU部署仍受硬件引擎未充分利用和加速器不兼容算子的限制,导致执行碎片化和每瓦吞吐量降低。本文提出了异构帧调度(H-FraDS),一种用于在NVIDIA边缘GPU上进行Transformer推理的硬件感知帧调度方法。H-FraDS使用固定调度比率在GPU和双深度学习加速器(DLA)核心之间路由帧,以提高延迟和功率限制下的利用率。为实现调度,通过重塑张量、用tanh近似误差函数(ERF)以及用有界tanh替换层归一化,使不兼容的Transformer组件适用于DLA执行。适配后的模型保持92%的F1分数,仅比原始模型降低2%。光流加速器(OFA)进一步用于推理侧光流估计。使用Swin Transformer进行自动驾驶感知,H-FraDS平衡调度(1:2)实现了125.93 FPS,比单独的适配DLA执行加速2.36倍,4.0 FPS/W,DLA延迟约24 ms,满足30 FPS实时操作;GPU-DLA-OFA情况实现了2.02倍的DLA吞吐量加速。
英文摘要
Physical AI systems, such as autonomous vehicles and intelligent machines, require transformer-based perception models that satisfy stringent edge latency and energy constraints. However, heterogeneous edge-GPU deployment remains limited by underutilized hardware engines and accelerator-incompatible operators, causing fragmented execution and lower throughput per watt. This paper presents Heterogeneous Frame Dispatch Scheduling (H-FraDS), a hardware-aware frame scheduling methodology for transformer inference on a recent NVIDIA edge GPU. H-FraDS routes frames across the GPU and dual deep learning accelerator (DLA) cores using fixed dispatch ratios to improve utilization under latency and power constraints. To enable scheduling, incompatible transformer components are adapted for DLA execution by reshaping tensors, approximating error function (ERF) with tanh, and replacing layer normalization with bounded tanh. The adapted model maintains a 92% F1 score, with only a 2% reduction from the original. Optical flow accelerator (OFA) is further used for inference-side optical-flow estimation. To the best of the authors' knowledge, prior work has not addressed these combined issues. Using Swin Transformer for autonomous-driving perception, H-FraDS Balanced Dispatch (1:2) achieves 125.93 FPS, a 2.36x speedup over standalone adapted-DLA execution, 4.0 FPS/W, and approximately 24 ms DLA latency, satisfying 30 FPS real-time operation; the GPU-DLA-OFA case achieves a 2.02x DLA throughput speedup.
Comments14 pages, 15 figures, This work has been submitted to IEEE for possible publication