arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在基础模型时代重新思考用于车辆重识别的多分支和跨骨干融合

Rethinking Multi-Branch and Cross-Backbone Fusion for Vehicle Re-Identification under Foundation-Model Pretraining

Yu Wang, Hongyu Yang

arXiv 2607.22068首次发表:更新:

发表机构

Huahuan (Yunnan) Technology Co., Ltd.(华环(云南)科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究在基础模型时代车辆重识别中多分支和跨骨干融合方法,以单个DINOv3预训练ConvNeXt为强基线,经无训练重排提升性能,通过实验评估增加表示多样性的效果,发现改进单骨干加检索重排比增加架构复杂度更有效。

AI 中文摘要

多分支架构和CNN-Transformer融合长期以来被视为通过组合互补表示来改进车辆重识别(Re-ID)的有效方法。在这项工作中,我们通过全面的实证研究在基础模型时代重新审视这一假设。一个经过调优的配方训练的单个DINOv3预训练ConvNeXt仅使用视觉线索在VeRi-Wild Small上达到88.19 mAP,在VeRi-Wild Large上达到77.47 mAP,与最强的协议验证的依赖元数据的多分支基线相匹配。应用无训练重排进一步将性能分别提高到92.38和83.68 mAP。使用这个强大的基线以及检索级分支诊断,我们评估增加表示多样性是否仍能带来可衡量的收益。在两个基准上,在共享骨干上连接多个分支改变最佳单分支性能不到1 mAP点,同时将嵌入维度增加4倍,并且得到的表示有效秩接近原始特征维度。我们进一步使用非对称冻结锚定策略研究跨骨干融合以组合ConvNeXt和视觉Transformer表示。尽管有这些有利条件,Transformer分支始终比ConvNeXt骨干低13 - 15 mAP,配对的逐查询自举分析估计观察到的最大融合增益仅为+0.11 mAP(95%置信区间)。我们的结果表明,在评估的设置下,改进单个强大的基础模型骨干以及检索阶段重排比通过额外分支或异构骨干增加架构复杂性更有效。我们将结论限制在单种子训练和一类基础模型,并讨论这些观察结果可能不成立的条件。

英文摘要

Multi-branch architectures and CNN-Transformer fusion are widely believed to improve vehicle re-identification (Re-ID) by combining complementary representations. We revisit this for a DINOv3-pretrained backbone. A single DINOv3-pretrained ConvNeXt with a tuned recipe reaches 88.19 mAP on VeRi-Wild Small and 77.47 on Large from visual cues alone, within the combined evaluation and optimization noise of the strongest protocol-verified metadata-dependent multi-branch baseline, and 92.38/83.68 with training-free re-ranking. Using this baseline and retrieval-level branch diagnostics, we ask whether representational diversity still pays at this scale. In our runs, it does not. Across both benchmarks and every converged configuration, concatenating multiple heads over a shared backbone moves the best single head by under one mAP point in either direction while costing four times the embedding dimension; 99.7% of the concatenation's variance lies in 512 principal components, so the heads not only duplicate one another but each occupies a quarter of its nominal 2048 dimensions. Pushing diversity to its architectural limit, CNN versus Transformer, we grant fusion every advantage through an asymmetric frozen-anchor scheme. Every Transformer configuration still lands at least 13 mAP below the ConvNeXt backbone (13-15 for the two strongest, up to 46 for the weakest), and a paired per-query bootstrap bounds the fusion gain at +0.11 mAP (95% CI) even for the most favourable snapshot we obtained. One strong backbone with the right recipe and re-ranking is the efficiency frontier. All results use single-seed training and one foundation-model family; differences of this size are therefore reported as bounds rather than orderings, and we list falsifiers.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑