REACT:用于实时事件驱动时间感知的全脉冲状态空间模型
REACT: A Fully Spiking State-Space Model for Real-Time Event-Driven Temporal Perception
浏览论文内容
中文总结 AI 辅助
提出全脉冲状态空间模型REACT,逐事件处理原始事件流,实现低延迟时间感知,在TTC估计中达到9.59%误差和4.6毫秒延迟,支持零样本迁移与INT8量化。
中文摘要 AI 辅助
在动态环境中运行的机器人系统需要随着传入的感官流连续演化的视觉感知。事件相机提供微秒级的时间分辨率和异步传感,但大多数基于学习的方法将事件累积到帧或时间箱中,引入了可能限制快速反应的积分延迟。在这里,我们提出了REACT,一种用于事件驱动时间感知的全脉冲状态空间模型,它逐事件处理原始事件,无需时间累积。REACT使用复值脉冲神经元C-SiLIF,其连续时间动力学由物理事件间间隔驱动,使其内部状态能够以单个事件的时间分辨率演化。我们在手势识别和从全场事件流中估计碰撞时间(TTC)上评估了REACT,无需目标边界框或定位输入。在EvTTC上,REACT实现了9.59%的相对TTC误差,端到端推理延迟为4.6毫秒,与最佳学习方法相差0.15个百分点,且无需目标先验。在数据集的平均接近速度下,该延迟仅对应4厘米的车辆运动,而最快的竞争学习方法对应1米。REACT进一步支持任意时间TTC预测、对不同驾驶序列的零样本迁移以及INT8量化,将每32,768个事件的估计能耗从18.5毫焦耳降低到2.8毫焦耳。这些结果表明,事件驱动的脉冲状态空间动力学可以为反应式机器人系统提供低延迟、持续更新的时间感知。
英文摘要
Robotic systems operating in dynamic environments require visual perception that evolves continuously with the incoming sensory stream. Event cameras provide microsecond temporal resolution and asynchronous sensing, but most learning-based methods accumulate events into frames or temporal bins, introducing an integration delay that can limit fast reaction. Here we propose REACT, a fully spiking state-space model for event-driven temporal perception that processes raw events one by one, without temporal accumulation. REACT uses a complex-valued spiking neuron, C-SiLIF, whose continuous-time dynamics are driven by the physical inter-event interval, allowing its internal state to evolve at the temporal resolution of individual events. We evaluate REACT on gesture recognition and time-to-collision (TTC) estimation from full-field event streams, without a target bounding box or localization input. On EvTTC, REACT achieves a 9.59% relative TTC error with 4.6 ms end-to-end inference latency, within 0.15 percentage points of the best learned method while requiring no target prior. At the dataset's mean approach speed, this latency corresponds to only 4 cm of vehicle motion, compared with 1 m for the fastest competing learned method. REACT further supports anytime TTC prediction, zero-shot transfer to a different driving sequence, and INT8 quantization, reducing the estimated energy consumption from 18.5 to 2.8 mJ per 32,768 events. These results show that event-driven spiking state-space dynamics can provide low-latency, continuously updated temporal perception for reactive robotic systems.
发表机构
- CerCo, CNRS UMR5549(CerCo,法国国家科学研究中心UMR5549)
- Université de Toulouse(图卢兹大学)
- IPAL, CNRS IRL, Singapore(IPAL,法国国家科学研究中心国际研究实验室,新加坡)
- ETIS, CY Cergy Paris Université(ETIS,CY塞尔吉巴黎大学)
- CNRS, ENSEA(法国国家科学研究中心,ENSEA)
机构由 AI 辅助整理,请以论文原文为准。