arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Paths:面向RGB-Event视频行人重识别的感知提示的时空Transformer与分层多模态融合

Paths: Prompt-aware Spatio-temporal Transformer with Hierarchical Multi-modal Fusion for RGB-Event Video Person Re-Identification

Yakun Huo, Yingquan Wang, Yangyang Liu, Tianyu Yan, Yunzhi Zhuge, Pingping Zhang, Huchuan Lu

arXiv 2608.13092首次发表:更新:

发表机构

Dalian University of Technology(大连理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有RE-VReID方法时空建模解耦、多模态融合不足的问题,提出含MAB、PST、HMF模块的Paths框架,在三个基准上验证了有效性。

AI 中文摘要

RGB-Event视频行人重识别(RE-VReID)旨在利用互补的RGB视频与事件流,在非重叠摄像头间检索特定行人。然而现有方法常将空间与时间建模解耦,限制了二者交互;且全局级RGB-事件融合无法充分挖掘细粒度判别线索。为解决这些问题,本文提出Paths,这是一个具备时空建模与分层多模态融合的RE-VReID统一框架。具体而言,首先设计记忆增强骨干网络(MAB)以维护模态特定的身份原型,实现稳定的模态内表示学习;接着提出感知提示的时空Transformer(PST),在统一Transformer内联合建模空间与时间线索;最后引入分层多模态融合(HMF),在全局与局部层面融合RGB与事件特征。借助这些模块,该框架可学习用于RE-VReID的鲁棒且具判别性的表示。在EvReID、MARS、iLIDS-VID三个公开RE-VReID基准上开展的大量实验,验证了所提方法的有效性,代码可在指定URL获取。

英文摘要

RGB-Event Video Person Re-Identification (RE-VReID) aims to retrieve specific person across non-overlapping cameras with complementary RGB videos and event streams. However, existing methods often decouple spatial and temporal modeling, which limits their interaction. In addition, global-level RGB-Event fusion fails to fully exploit fine-grained discriminative cues. To address these issues, we propose Paths, a unified framework with spatio-temporal modeling and hierarchical multi-modal fusion for RE-VReID. Specifically, we first design a Memory-Augmented Backbone (MAB) to maintain modality-specific identity prototypes for stable intra-modal representation learning. Then, we propose a Prompt-aware Spatio-temporal Transformer (PST) to jointly model spatial and temporal cues within a unified Transformer. Finally, we introduce a Hierarchical Multi-modal Fusion (HMF) to integrate RGB and event features at global and local levels. With these modules, our framework can learn robust and discriminative representations for RE-VReID. Extensive experiments on three public RE-VReID benchmarks including EvReID, MARS and iLIDS-VID, demonstrate the effectiveness of our proposed method. The code is available at https://github.com/Reflection0427/Paths.

CommentsAccepted by ACM MM2026. More modifications may be performed

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑