发表机构
Istituto Italiano di Tecnologia(意大利理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种无需训练、基于显著性的多尺度注意力模型,直接处理低分辨率事件流以选择ROI,在Prophesee数据集上以毫秒级分辨率实现高达70.8%的多类目标检测准确率,为边缘神经形态视觉提供高效前端。
AI 中文摘要
神经形态视觉系统在带宽、内存和能量方面受到严格限制,尤其是在边缘计算场景中,这促使了早期数据缩减和选择性处理机制的发展。在本工作中,我们研究了一种多尺度、无需训练、基于显著性的自底向上视觉注意力模型,该模型直接处理低分辨率的事件输入,并从视觉场景中选择感兴趣区域(ROI)。该模型在应用于传入事件流的多个降采样因子下进行评估,输入分辨率相对于全分辨率最多可降低256倍。性能在Prophesee Automotive数据集上进行了评估,该数据集是最大的公开可用事件数据集,展示了在真实世界用例中不同尺度下的稳健ROI选择。所提出的方法能够检测属于多个对象类别的ROI,包括各种车辆类型、行人、交通灯和交通标志,准确率高达70.8%,同时以毫秒级时间分辨率运行,比数据集提供的真值时间分辨率精细16倍。这些结果凸显了将早期事件降采样与基于显著性的注意力相结合,作为高效边缘神经形态视觉系统有效前端的潜力。
英文摘要
Neuromorphic vision systems operate under strict constraints on bandwidth, memory, and energy, particularly at the edge, motivating early mechanisms for data reduction and selective processing. In this work, we investigate a multi-scale training-free, saliency-based, bottom-up visual attention model that operates directly on low-resolution event-based input and selects Regions of Interest (ROI) from the visual scene. The model is evaluated across multiple downscaling factors applied to the incoming event stream, with input resolutions reduced by up to 256x relative to full resolution. Performance is assessed on the Prophesee Automotive dataset, the largest publicly available event-based dataset, demonstrating robust ROI selection across different scales on a real-world use-case. The proposed approach is capable of detecting ROIs belonging to multiple object classes, including various vehicle types, pedestrians, traffic lights, and traffic signs, with accuracy up to 70.8%, while operating at millisecond temporal resolution, 16x finer than the temporal resolution provided by the dataset ground truth. These results highlight the potential of combining early event downscaling with saliency-based attention as an effective front-end for efficient edge neuromorphic vision systems.