arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估用于车辆属性识别的2D与3D感知视觉基础模型

Evaluating 2D and 3D-Aware Vision Foundation Models for Vehicle Attribute Recognition

Alexandre V. Delazeri, Gabriel E. Lima, Eduil Nascimento, Rayson Laroca, David Menotti

arXiv 2608.29929首次发表:更新:

发表机构

Federal University of Paraná; Paraná Military Police; Pontifical Catholic University of Paraná(巴拉那联邦大学; 巴拉那宪兵队; 巴拉那天主教大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过对14种2D与3D感知视觉基础模型的实证基准测试,发现DINOv3等2D自监督模型在车辆属性识别细粒度任务中优于多数3D感知模型,仅Depth Anything v2在车辆类型视角鲁棒性上更优,为混合方法提供了依据。

AI 中文摘要

车辆属性识别是智能交通系统中的一项重要任务,尤其在自动车牌识别(ALPR)不可用或不可靠时更为关键。尽管视觉基础模型已展现出跨领域的强迁移能力,但它们在细粒度车辆分类中的有效性仍未得到充分探索。此外,鉴于车辆具有固有的三维结构,目前尚不清楚新兴的3D感知基础模型是否比标准2D架构更具优势。本文对14种最先进的2D和3D感知视觉基础模型开展了实证基准测试,采用具有挑战性的真实世界UFPR-VeSV数据集,将这些模型作为冻结特征提取器,通过线性探测评估其在车辆类型、品牌和型号识别上的性能。我们进一步在少样本学习和分布外(OOD)域偏移场景下对表现最佳的模型进行了压力测试。结果显示,标准2D自监督模型,特别是DINOv3,在细粒度任务中大幅优于3D感知模型,在车辆品牌和型号识别上的宏观准确率超过93%;但3D感知的Depth Anything v2在车辆类型分类中对视角变化的鲁棒性更强。这些发现为结合2D与3D先验的混合方法提供了依据,以实现鲁棒的车辆识别。我们的代码已在该URL公开可用。

英文摘要

Vehicle attribute recognition is an important task in intelligent transportation systems, particularly when Automatic License Plate Recognition (ALPR) is unavailable or unreliable. Although vision foundation models have shown strong transferability across domains, their effectiveness for fine-grained vehicle classification remains underexplored. Moreover, given the inherently three-dimensional structure of vehicles, it is unclear whether emerging 3D-aware foundation models offer advantages over standard 2D architectures. This paper presents an empirical benchmark of 14 state-of-the-art 2D and 3D-aware vision foundation models. Using the challenging real-world UFPR-VeSV dataset, we evaluate these models as frozen feature extractors via linear probing for vehicle type, make, and model recognition. We further stress-test the best-performing models under few-shot learning and Out-of-Distribution (OOD) domain shifts. Our results show that standard 2D self-supervised models, particularly DINOv3, substantially outperform 3D-aware models in fine-grained tasks, achieving over 93% Macro-Accuracy for make and model recognition. However, the 3D-aware Depth Anything v2 exhibits stronger invariance to viewing angles in vehicle type classification. These findings motivate hybrid approaches that combine 2D and 3D priors for robust vehicle recognition. Our code is publicly available at https://github.com/UFPR-IPASPPR/3D-Vision-Benchmark/.

CommentsAccepted for presentation at the 2026 Conference on Graphics, Patterns and Images (SIBGRAPI)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑