AI 中文总结
本文提出无监督脉冲特征学习方法,在DVS-PedX基准上以90.3%准确率接近监督性能,表明标签代价小且弱点源于读出协议。
AI 中文摘要
事件相机非常适合行人穿越检测,而脉冲神经网络(SNNs)能够原生处理其输出,但当前的SNN检测器使用监督反向传播进行训练,因此需要昂贵的帧级穿越标签。我们研究了在特征学习中无标签情况下的穿越检测,并在最新的DVS-PedX行人穿越基准上报告了首个无监督结果。一个使用赢者通吃脉冲时序依赖可塑性的单脉冲层从无标签的事件帧补丁中学习字典;帧通过余弦相似度编码到学习到的滤波器,并进行空间池化,然后由线性分类器读出,这是唯一的监督组件。在24,454帧的测试集上,该方法达到了90.3%的准确率和0.936的接收者操作特征曲线下面积(AUROC),而一个在相同帧上端到端训练的监督脉冲网络达到了92.0%和0.943;在恶劣天气下,AUROC为0.913。该结果对可塑性规则的选择不敏感,但强烈依赖于读出协议:使用经典的神经元分配读出,同一网络仅达到0.70的AUROC。在基准的真实转换部分,重新拟合的读出将性能从随机水平(0.53)提升到0.67的AUROC,处于已发表的监督范围内。这些发现表明,在该基准上,从特征学习中移除标签的准确率代价很小,并且所报道的无监督脉冲网络的弱点可能归因于读出协议而非学习规则。
英文摘要
Event cameras are well suited to pedestrian crossing detection, and spiking neural networks (SNNs) can process their output natively, but current SNN detectors are trained with supervised backpropagation and therefore require costly frame-level crossing labels. We investigate crossing detection with no labels in feature learning and report the first unsupervised results on the recent DVS-PedX pedestrian-crossing benchmark. A single spiking layer trained with winner-take-all spike-timing-dependent plasticity learns a dictionary from unlabelled event-frame patches; frames are encoded by cosine similarity to the learned filters with spatial pooling and read out by a linear classifier, the only supervised component. On the 24,454-frame test split the method attains 90.3% accuracy and an area under the receiver operating characteristic curve (AUROC) of 0.936, compared with 92.0% and 0.943 for a supervised spiking network trained end-to-end on the same frames; under adverse weather the AUROC is 0.913. The result is insensitive to the choice of plasticity rule but depends strongly on the readout protocol: with the classical neuron-assignment readout the same network attains only 0.70 AUROC. On the benchmark's real converted portion, a readout refit lifts performance from chance (0.53) to 0.67 AUROC, within the published supervised range. These findings indicate that the accuracy cost of removing labels from feature learning is small on this benchmark, and that reported weaknesses of unsupervised spiking networks may be attributable to the readout protocol rather than to the learning rule.
Comments18 pages, 4 figures, 3 tables