发表机构
NKIARI; VCIP, CS, Nankai University; AAIS, Nankai University(鹏城实验室; 南开大学视觉计算与图像处理研究室; 南开大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对RGB-事件解析中统一主干应用的挑战,提出Evita统一主干,嵌入几何视差校正等协同学习模块,构建N-ImageNetV2及预训练协议,经多基准评估,其实现新指标,在实时多模态上精度-延迟权衡表现卓越。
AI 中文摘要
融合标准RGB帧与异步事件流已成为在退化环境中进行稳健感知的决定性范式。尽管统一主干最近在多模态视觉中受到关注,但将它们应用于RGB-事件领域仍具有根本挑战性。现有架构要么采用解耦双编码器,增加计算开销,要么采用通用统一设计,无法解决密集强度网格和稀疏运动尖峰之间极端表示差异下的隐式几何视差和跨光谱混叠。为克服这些瓶颈,我们提出了Evita,首个专为密集RGB-事件解析设计的统一主干。为实现深度模态协同,Evita在每个编码器层直接嵌入一组内在协同学习模块,包括用于自适应空间对齐的几何视差校正、仅在复频域进行纹理转移的谐波光谱共振以及用于事件驱动非对称注意力的瞬态全局路由。为保证针对空间错位的稳健特征提取并将表示与特定事件编码解耦,我们构建了N-ImageNetV2并采用随机事件表示混合预训练协议,使网络能在下游任务中无缝适应任意事件格式。在DELIVER、DDD17和DSEC基准上的广泛评估证实,Evita建立了新的最先进指标,同时在实时多模态方面实现了卓越的精度-延迟权衡。
英文摘要
Fusing standard RGB frames with asynchronous event streams has emerged as a definitive paradigm for robust perception in degraded environments. Although unified backbones have recently gained traction in multi-modal vision, adapting them to the RGB-Event domain remains fundamentally challenging. Existing architectures either resort to decoupled dual encoders that double computational overhead, or adopt generic unified designs that fail to resolve implicit geometric parallax and cross-spectral aliasing under the extreme representational divide between dense intensity grids and sparse kinematic spikes. To transcend these bottlenecks, we present Evita, the first unified backbone specifically engineered for dedicated dense RGB-Event parsing. To achieve profound modal synergy, Evita explicitly embeds a suite of intrinsic co-learning modules directly into every encoder layer. Specifically, it features Geometric Parallax Rectification for adaptive spatial alignment, Harmonic Spectral Resonance for texture transfer exclusively in the complex frequency domain, and Transient Global Routing for event-driven asymmetric attention. To guarantee robust feature extraction against spatial misalignments and decouple representations from specific event encodings, we construct N-ImageNetV2 alongside a stochastic event representation mixing pretraining protocol, empowering the network to seamlessly accommodate arbitrary event formats in downstream tasks. Extensive evaluations across the DELIVER, DDD17, and DSEC benchmarks confirm that Evita establishes new state-of-the-art metrics while delivering a superior accuracy-latency trade-off for real-time multimodal perception.The code are publicly available at: https://github.com/chaineypung/Evita.