发表机构
Khalifa University(哈利法大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种软硬件协同方法,利用解析或CNN运动矢量传播实现高效视频目标检测,在降低延迟和能耗的同时保持准确率,并验证了联合设计时间模型与执行管道的价值。
AI 中文摘要
连续视频分析要求在嵌入式功耗预算内以低延迟实现准确定位。本文提出了一种软硬件协同设计方法,在检测器调用之间重用编解码器运动矢量(MVs)。两种备选模型支持平移和尺度变化:解析运动矢量传播(Analytical-MV)和使用卷积神经网络(CNN)的学习式传播(CNN-MV)。学习模型采用卷积操作和独立目标更新,适合在边缘图形处理单元(GPU)上并行执行。Analytical-MV结合了0.909的调和平均精确率-召回率分数(F1)、9.03毫秒的平均端到端延迟和每帧0.177焦耳的能量消耗,在所评估的配置中实现了最低延迟和能量消耗。相对于每帧检测,它将平均延迟降低了25.9%,每帧能量降低了36.4%。CNN-MV提供了不同的权衡:其最快配置将召回率从Analytical-MV的0.871提高到0.890,并将平均功率从19.64瓦降至17.32瓦,同时实现了18.42毫秒的延迟和每帧0.319焦耳的能量消耗。因此,当召回率或运行功率比最小延迟和能量更重要时,它很有用。在深度学习加速器(DLA)上执行进一步降低了相对于GPU执行的时间平均GPU利用率。主机处理优化显著改善了延迟和能量,证明了联合设计时间模型及其执行管道的价值。
英文摘要
Continuous video analytics requires accurate localization at low latency within embedded power budgets. This paper presents a hardware-software design methodology that reuses codec motion vectors (MVs) between detector invocations. Two alternative models support translation and scale changes: analytical motion-vector propagation (Analytical-MV) and learned propagation using a convolutional neural network (CNN) (CNN-MV). The learned model uses convolutional operations and independent object updates suited to parallel execution on an edge graphics processing unit (GPU). Analytical-MV combines a harmonic-mean precision-recall score (F1) of 0.909 with a mean end-to-end latency of 9.03 ms and an energy consumption of 0.177 J per frame, yielding the lowest latency and energy among the evaluated configurations. Relative to detection on every frame, it reduces mean latency by 25.9% and energy per frame by 36.4%. CNN-MV offers a different trade-off: its fastest configuration raises recall from 0.871 for Analytical-MV to 0.890 and lowers mean power from 19.64 to 17.32 W, while achieving a latency of 18.42 ms and an energy consumption of 0.319 J per frame. It is therefore useful when recall or operating power is more important than minimum latency and energy. Execution on a deep learning accelerator (DLA) further reduces time-averaged GPU utilization relative to GPU execution. Host-processing optimization substantially improves both latency and energy, demonstrating the value of jointly designing temporal models and their execution pipelines.
Comments11 pages, 4 figures, This work has been submitted to IEEE for possible publication