发表机构
Delft University of Technology; KU Leuven; University of Zurich; University of Pennsylvania(代尔夫特理工大学; 鲁汶大学; 苏黎世大学; 宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对边缘高分辨率视觉的低延迟需求,研究人员推出首款EV-GNN处理器ETHEREAL,通过邻域并行样条卷积引擎等实现25.6μs推理延迟,在DAGr-GNN工作负载等上完成验证。
AI 中文摘要
动态视觉传感器(DVS)是边缘视觉应用实现低延迟、亚毫秒目标的诱人候选,因为它能生成具有微秒级时间分辨率的事件。然而,使用DVS前端也需要新颖的算法/硬件后端,能够高效处理稀疏时空事件流。虽然事件驱动图神经网络(EV-GNNs)在算法层面已成为一种既准确又高效的解决方案,但迄今为止还没有专用硬件能够有效支持其对密集规则计算操作和稀疏不规则内存访问的混合需求。因此,我们推出ETHEREAL,这是首款EV-GNN处理器芯片,通过邻域并行样条卷积引擎与拆分2D/3D内存层次结构相结合,引入了一种新颖的时空事件缓存机制,从而弥合了这一差距。测量结果表明,在最先进的DAGr-GNN工作负载和VGA分辨率(640×480像素)的DSEC数据集上,每端到端事件推理的延迟为25.6μs,能耗为1.6μJ。
英文摘要
Dynamic vision sensors (DVS) are enticing candidates to reach the low-latency, sub-ms target of edge-vision applications, as they generate events with a $μ$s-level time resolution. However, using DVS front ends also calls for novel algorithm/hardware back ends capable of efficiently handling streams of sparse spatiotemporal events. While event-driven graph neural networks (EV-GNNs) have emerged as a solution on the algorithmic side that is both accurate and efficient, there is no dedicated hardware to date capable of efficiently supporting their mixed requirements of dense-regular compute operations and sparse-irregular memory accesses. We therefore introduce ETHEREAL, the first EV-GNN processor chip, capable of bridging this gap by means of a neighbor-parallel spline-convolution engine combined with a split-2D/3D memory hierarchy that introduces a novel spatiotemporal event-caching mechanism. Measurement results demonstrate a 25.6$μ$s latency and a 1.6$μ$J energy per end-to-end event-wise inference on the state-of-the art DAGr-GNN workload and VGA-resolution (640x480 pixels) DSEC dataset.
CommentsThis work has been submitted to the IEEE JSSC for possible publication