一种具有时空缓存和样条卷积的25-$\mu$s/inf事件驱动图神经网络处理器,用于边缘超低延迟AI
A 25-$μ$s/inf Event-driven Graph Neural Network Processor with Spatiotemporal Caching and Spline Convolution for Ultra-low-latency AI at the Edge
- Delft University of Technology (TU Delft)(代尔夫特理工大学)
- KU Leuven(鲁汶大学)
- University of Zürich (UZH)(苏黎世大学)
- University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对DVS事件流超低延迟边缘AI,提出首个可扩展至640×480分辨率的EV-GNN加速器ETHEREAL,采用样条卷积与时空缓存,实现25.6μs延迟和1.7μJ/事件能耗。
AI中文摘要:
动态视觉传感器(DVS)相机以逐像素为基础生成事件,时间分辨率达微秒级,与标准帧基视觉相比,需要新的算法-硬件协同设计方法。虽然事件驱动图神经网络(EV-GNNs)成为一种有前景的算法解决方案,但它们通过混合密集规则计算操作和稀疏不规则内存访问带来了新的硬件挑战。我们提出了ETHEREAL,这是首个可扩展到640$\times$480分辨率的EV-GNN加速器,得益于邻域并行样条卷积引擎和具有新颖感兴趣区域时空缓存机制的2D/3D分割内存层次结构。测量结果表明,在最先进的工作负载上,端到端推理延迟为25.6$\mu$s,每事件能量为1.7$\mu$J。
英文摘要:
Dynamic-vision-sensor (DVS) cameras generate events on a per-pixel basis with a $μ$s-level temporal resolution, calling for new algorithm-hardware co-design approaches compared to standard frame-based vision. While event-driven graph neural networks (EV-GNNs) emerge as a promising algorithmic solution, they raise new HW challenges by mixing dense-regular compute operations and sparse-irregular memory accesses. We present ETHEREAL, the first EV-GNN accelerator that scales to 640$\times$480 resolutions, thanks to a neighbor-parallel spline convolution engine and a 2D/3D-split memory hierarchy with a novel region-of-interest spatiotemporal caching mechanism. Measurement results demonstrate end-to-end inference with 25.6$μ$s latency and 1.7$μ$J energy per event on state-of-the-art workloads