arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22500cs.CV

事件-帧融合用于帧间分割:事件引导的运动

Event-Frame Fusion for Inter-Frame Segmentation via Event-Guided Motion

Dalia Hareb, Jean Martinet, Benoit Miramond, Elisabetta Chicca

首次发表
浏览论文内容

中文总结 AI 辅助

针对传统相机帧间信息丢失问题,提出结合帧与事件相机的混合架构,利用紧凑SNN估计运动并插值帧间分割,实现高达500Hz分割率且低能耗,提升鲁棒性。

中文摘要 AI 辅助

自主导航需要精确且高效的语义分割,然而现有的基于帧的方法仍受限于运动模糊、眩光、延迟以及传统相机较低的时间分辨率(20-30 FPS),这导致帧间信息丢失。事件相机作为一种替代传感模态出现,它以异步方式捕捉强度变化,具有高时间分辨率、高动态范围和稀疏输出。然而,基于事件的算法在准确性上仍不及基于帧的算法,因为大多数分割方法是为密集帧数据设计的。为克服这些限制,我们提出了一种混合视觉架构,结合了传统的基于帧的相机和基于事件的相机。该系统集成了两个互补组件:(1)一个紧凑的脉冲神经网络(SNN),具有42k参数,用于运动估计;(2)一个轻量级的事件驱动SNN,具有0.84M参数,用于基于帧的语义分割,该网络在帧间插值运动以细化分割结果。通过预测帧间分割,该框架实现了高达500 Hz的分割率,每次推理能耗低于1.87 mJ,同时以高达200 Hz的频率保持实时GPU执行。此外,我们的方法补偿了受模糊或过度曝光影响的帧中的信息损失,从而在挑战性条件下实现更鲁棒的感知。

英文摘要

Autonomous navigation requires precise and efficient semantic segmentation, yet existing frame-based approaches remain limited by motion blur, glare, latency, and the low temporal resolution (20-30 FPS) of conventional cameras, which leads to information loss between frames. Event cameras have emerged as an alternative sensing modality, capturing intensity changes asynchronously with high temporal resolution, high dynamic range, and sparse outputs. However, event-based algorithms still fall short of frame-based ones in accuracy, as most segmentation methods are designed for dense frame data. To overcome these limitations, we propose a hybrid vision architecture that combines conventional frame-based and event-based cameras. The system integrates two complementary components: (1) a compact Spiking Neural Network (SNN) with 42k parameters for motion estimation, and (2) a lightweight event-driven SNN with 0.84M parameters for frame-based semantic segmentation, which interpolates motion between frames to refine segmentation results. By predicting inter-frame segmentations, the framework achieves segmentation rates of up to 500 Hz with an energy consumption below 1.87 mJ per inference, while maintaining real-time GPU execution at frequencies up to 200 Hz. Additionally, our approach compensates for information loss in frames affected by blur or overexposure, enabling more robust perception in challenging conditions.

发表机构

  • I3S, Côte d’Azur University, CNRS(I3S,蔚蓝海岸大学,法国国家科学研究中心)
  • LEAT, Côte d’Azur University(LEAT,蔚蓝海岸大学)
  • Zernike Institute, Groningen University(泽尼克研究所,格罗宁根大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑