arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13328cs.CVcs.AI

基于事件视觉的卷积脉冲神经网络与时间增强用于行人过街意图分类

Pedestrian Crossing Intent Classification From Event-Based Vision Using Convolutional Spiking Neural Networks With Temporal Augmentation

Henok Teklu, Mustafa Sakhai, Maciej Wielgosz, Matej Mertik

首次发表
浏览论文内容

中文总结 AI 辅助

针对自动驾驶中行人过街意图预测,提出基于事件视觉的卷积脉冲神经网络,利用v2e和CARLA生成DVS数据增强训练,在JAAD上达95.83%准确率,实现高效实时分类。

中文摘要 AI 辅助

预测行人是否会过马路对于自动驾驶汽车的安全至关重要,这要求在运动模糊、高动态范围和类别不平衡等具有挑战性的条件下进行实时推理。传统的基于帧的深度网络以固定帧率处理冗余的RGB数据,限制了其时间分辨率和能量效率。在本工作中,我们提出了一种端到端流水线,该流水线(i)使用v2e模拟器将来自联合注意力自动驾驶(JAAD)数据集的真实驾驶视频转换为合成动态视觉传感器(DVS)事件流,(ii)使用DVS-PedX数据集的CARLA模拟DVS序列在正常和恶劣天气条件下增强训练,(iii)训练一种新颖的卷积脉冲神经网络(Conv-SNN),并采用片段一致的DVS增强来将行人过街意图分类为二分类:过街或不过街。我们详细说明了所有架构决策、精确的泄漏积分发放神经元动力学与替代梯度学习、类别平衡损失公式、JAAD的6倍过采样以及70/15/15分层划分协议。训练后的模型在JAAD DVS测试集上达到95.83%的准确率和F1=0.9695,在正常CARLA DVS上达到97.79%的准确率和F1=0.9478,在恶劣天气CARLA DVS上达到94.78%的准确率和F1=0.8369,所有这些均来自一个在CPU上训练的1.07M参数架构。与先前基于帧的JAAD方法相比,我们的方法在原生处理稀疏时间表示的同时,缩小或超越了所报告的准确率。我们包含了对所有15个训练轮次的收敛行为、域迁移特性以及与代表性相关工作的定量比较的全面分析。

英文摘要

Anticipating whether a pedestrian will cross the road is safety-critical for autonomous vehicles, requiring real-time inference under challenging conditions including motion blur, high dynamic range, and class imbalance. Conventional frame-based deep networks process redundant RGB data at fixed frame rates, limiting their temporal resolution and energy efficiency. In this work we present an end-to-end pipeline that (i) converts real-world driving footage from the Joint Attention in Autonomous Driving (JAAD) dataset into synthetic dynamic vision sensor (DVS) event streams using the v2e simulator, (ii) augments training with the CARLA-simulated DVS sequences of the DVS-PedX dataset under both normal and adverse weather conditions, and (iii) trains a novel convolutional spiking neural network (Conv-SNN) with clip-consistent DVS augmentation to classify pedestrian crossing intent as binary: crossing or non-crossing. We detail all architectural decisions, the exact leaky-integrate-and-fire neuron dynamics with surrogate-gradient learning, the class-balanced loss formulation, JAAD oversampling at 6x, and a 70/15/15 stratified splitting protocol. The trained model achieves 95.83% accuracy and F1 = 0.9695 on the JAAD DVS test set, 97.79% accuracy and F1 = 0.9478 on normal CARLA DVS, and 94.78% accuracy and F1 = 0.8369 on adverse-weather CARLA DVS, all from a 1.07M-parameter architecture trained on CPU. Compared to prior frame-based approaches on JAAD, our method closes or surpasses the reported accuracy while operating natively on sparse temporal representations. We include a thorough analysis of the convergence behaviour across all 15 training epochs, domain transfer characteristics, and a quantitative comparison with representative related work.

发表机构

  • Alma Mater Europaea University(欧洲母校大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑