arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39492cs.CV

Front-to-Back:用于非对称跨视角车辆再识别的视觉-语言模型基准测试

Front-to-Back: Benchmarking Vision-Language Models for Asymmetric Cross-View Vehicle Re-Identification

Moseli Mots'oehli, Thulani Babeli

首次发表
浏览论文内容

中文总结 AI 辅助

针对跨视角车辆再识别难题,提出Front2Back-ReID基准,评估零样本视觉-语言模型与检索基线,发现VLM未稳定超越强检索,人类准确率显著更高。

中文摘要 AI 辅助

匹配同一车辆在前置摄像头和后置摄像头中的图像是困难的,因为摄像头之间没有共享视角,且车辆外观变化显著。我们提出了Front2Back-ReID基准,该基准包含来自南非20个录制序列的500个经人工验证的车辆交接案例。每个案例要求模型将前置摄像头图像中高亮显示的车辆与后置摄像头图像中至少三个候选车辆中的同一车辆进行匹配。我们评估了七个零样本视觉-语言模型、四个图像检索基线以及25名人类参与者。模型使用完整的前置RGB图像、裁剪的目标车辆以及二值轮廓进行测试。最强的VLM在目标裁剪上达到了76.6%的Rank-1准确率,而冻结的SigLIP2基线为74.0%;这一差异在统计上并不显著。人类参与者使用完整图像达到了94.0%的准确率,使用目标裁剪达到了92.2%的准确率。在我们的评估设置下,启用推理在所有三种输入条件下均提高了每个以两种模式评估的模型的准确率。我们还发现,视觉-语言模型在完整场景上的表现通常不如在目标裁剪上的表现。这些结果表明,通用的视觉-语言模型在前置到后置车辆匹配方面尚未持续优于强视觉检索,而人类仍然明显更可靠。

英文摘要

Matching the same vehicle across front and rear cameras is difficult because the cameras do not share a view and the vehicle's appearance changes substantially. We introduce Front2Back-ReID, a benchmark of 500 manually verified vehicle handovers from 20 recording sequences in South Africa. Each example asks a model to match a vehicle highlighted in a front-camera image to the same vehicle among at least three candidates in a later rear-camera image. We evaluate seven zero-shot vision-language models, four image-retrieval baselines, and 25 human participants. Models are tested using full front RGB images, cropped target vehicles, and binary silhouettes. The strongest VLM achieved 76.6 percent Rank-1 accuracy on target crops, compared with 74.0 percent for the frozen SigLIP2 baseline; this difference was not statistically clear. Human participants achieved 94.0 percent accuracy with full images and 92.2 percent with target crops. Under our evaluation setup, enabling reasoning improved accuracy across all three input conditions for every model evaluated in both modes. We also found that VLMs generally performed worse on full scenes than on target crops. These results show that general-purpose VLMs do not yet consistently outperform strong visual retrieval for front-to-rear vehicle matching, while humans remain substantially more reliable.

发表机构

  • MindForge AI

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑