arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VR3D:面向空地行人重识别的视角鲁棒3D表示学习

VR3D: View-Robust 3D Representation Learning for Aerial-Ground Person Re-Identification

Chao Ji, Shiyu Xuan, Zechao Li

arXiv 2608.02598首次发表:更新:

AI 中文总结

本文针对空地行人重识别的视角偏差问题,提出VR3D框架,通过视角鲁棒3D表示交互与可靠性感知融合技术,在三个基准数据集上实现了优于现有方法的性能,CARGO数据集Rank-1提升5.63%。

AI 中文摘要

空地行人重识别是一项极具挑战性的任务,跨平台视角变化会导致严重的遮挡和几何变形。现有方法仅在2D图像空间内学习视角不变表示,而剧烈的视角变化会使学习到的特征仍与视角偏差耦合。为解决该问题,本文提出VR3D——一种视角鲁棒3D表示学习框架,该框架将图像映射到统一的3D坐标空间以实现视角无关的特征交互。具体而言,本文引入视角鲁棒3D表示交互(View-Robust 3D Representation Interaction),利用从单张2D观测中提取的3D先验,将2D外观特征提升到规范3D空间;VR3I采用3D几何-语义注意力(3D Geometry-Semantic Attention),基于对应身体部位的3D空间位置建立2D图像块与3D体素之间的交互,有效将2D语义锚定在3D框架内。此外,由于视角变化和3D重建误差会导致不同样本的表示可靠性存在差异,本文引入可靠性感知融合(Reliability-Aware Fusion),该模块会估计样本特定的可靠性并自适应聚合多源表示。在三个基准数据集(CARGO、AG-ReID.v1和AG-ReID.v2)上进行的大量实验表明,VR3D的性能优于现有最新方法,例如其在CARGO数据集上的Rank-1指标提升了5.63%,本文的代码将公开。

英文摘要

Aerial-ground person re-identification is a challenging task due to cross-platform viewpoint variations, which cause severe occlusion and geometric deformation. Existing methods attempt to learn view-invariant representations exclusively within the 2D image space, where drastic viewpoint variations cause the learned features to remain coupled with viewpoint bias. To address this, we propose VR3D, a View-Robust 3D Representation Learning framework that maps images into a unified 3D coordinate space to achieve view-independent feature interaction. Specifically, we introduce View-Robust 3D Representation Interaction, which leverages 3D priors extracted from single 2D observations to lift 2D appearance features into a canonical 3D space. VR3I employs 3D Geometry-Semantic Attention to establish interactions between 2D patches and 3D voxels from corresponding body parts based on their 3D spatial locations, effectively grounding 2D semantics within a 3D framework. In addition, as the reliability of these representations varies across samples due to viewpoint changes and 3D reconstruction errors, we introduce Reliability-Aware Fusion, which estimates sample-specific reliability and adaptively aggregates the multi-source representations. Extensive experiments on three benchmark datasets (CARGO, AG-ReID.v1, and AG-ReID.v2) demonstrate that VR3D outperforms recent methods. For example, it achieves a 5.63% improvement in Rank-1 on CARGO. Our code will be released.

Comments12 pages, 10 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑