arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26919cs.ROcs.CV

拨快时钟:在事件驱动目标检测中提前预测以克服延迟

Bend the Clock: Predicting Ahead to Beat Latency in Event-Based Object Detection

Biswadeep Sen, Benoit R. Cottereau, Nicolas Cuperlier, Terence Sim

首次发表
浏览论文内容

中文总结 AI 辅助

针对事件检测器计算延迟导致预测过时的问题,提出因果可用性时间检测器ChronoFuse,通过轻量级跨时间融合提前预测输出时刻目标状态,在多个数据集上显著恢复延迟损失,最高达9.3倍性能提升。

中文摘要 AI 辅助

事件相机有望为高速机器人系统提供低延迟感知,在这些系统中,即使是短暂的延迟也可能导致检测结果在用于下游机器人决策时已经过时。然而,现代事件检测器在产生预测之前仍需要数十毫秒的计算时间。传统评估忽略了这一延迟,将预测与观测时间戳处的标注进行比较,即使场景在预测生成时可能已经发生变化。我们研究了事件驱动多目标检测中的这种观测-可用性不匹配问题,并表明最先进的事件检测器在按预测可用时间而非观测时间评估时性能显著下降。为解决这一问题,我们提出了ChronoFuse,一种因果可用性时间检测器,它预测输出可用时的目标状态,而非输入被观测时的状态。ChronoFuse在多尺度特征层级上进行因果跨时间融合,将当前表示与缓存的时间特征相结合,以在不使用未来观测的情况下暴露短期时间线索。该融合路径轻量级,仅增加0.17百万参数和0.84毫秒的平均端到端延迟开销。ChronoFuse在1Mpx驾驶数据上恢复了因延迟损失的71%的准确率,在FRED上快速无人机运动下恢复了90.8%,几乎恢复了零延迟性能。在EV-Flying的极端运动下,ChronoFuse达到20.95 sAP,而最强的标准事件检测器仅为2.25(9.3倍增益)。这些结果表明,对于在快速变化场景中运行的机器人,包括自动驾驶、敏捷飞行和机器人拦截,提前预测至关重要。

英文摘要

Event cameras promise low-latency perception for high-speed robotic systems, where even short delays can render detections stale by the time they inform downstream robotic decisions. Yet modern event detectors still require tens of milliseconds of computation before their predictions become available. Conventional evaluation ignores this delay by comparing predictions with annotations at the observation timestamp, even though the scene may have changed by the time those predictions are produced. We study this observation-availability mismatch in event-based multi-object detection and show that state-of-the-art event detectors degrade substantially when evaluated at prediction availability rather than observation time. To address this, we introduce ChronoFuse, a causal availability-time detector that predicts object states for when its output becomes available rather than for when its input was observed. ChronoFuse performs causal cross-time fusion over a multi-scale feature hierarchy, combining current representations with cached temporal features to expose short-term temporal cues without using future observations. The fusion pathway is lightweight, adding only 0.17 million parameters and 0.84 ms of mean end-to-end latency overhead. ChronoFuse recovers 71% of the accuracy lost to latency on 1Mpx driving data and 90.8% under rapid drone motion on FRED, nearly restoring zero-delay performance. Under the extreme motion of EV-Flying, ChronoFuse reaches 20.95 sAP, compared with 2.25 for the strongest standard event detector (9.3x gain). These results show that predicting ahead can be critical for robots operating in fast-changing scenes, including autonomous driving, agile flight, and robotic interception.

发表机构

  • National University of Singapore(新加坡国立大学)
  • IPAL CNRS IRL 2955(法国国家科研中心IPAL联合研究实验室2955)
  • CerCo, CNRS UMR 5549, Université de Toulouse(法国国家科研中心CerCo实验室UMR 5549,图卢兹大学)
  • ETIS UMR8051, CY Cergy Paris Université, ENSEA, CNRS(法国国家科研中心ETIS实验室UMR8051,CY塞尔吉巴黎大学,ENSEA)

机构由 AI 辅助整理,请以论文原文为准。

↑