arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15563cs.CV

视觉场所识别中所有令牌都必不可少吗?高效推理的令牌约简实证研究

Are All Tokens Necessary for Visual Place Recognition? An Empirical Study of Token Reduction for Efficient Inference

Tong Jin, Yunpeng Liu, Shuyu Hu, Qinghua Zhang, Ruize Han, Song Wang, Feng Lu

首次发表
浏览论文内容

中文总结 AI 辅助

研究视觉场所识别中是否所有令牌都必要,提出首个令牌约简系统基准,在多模型和数据集上评估多种方法,从多视角研究令牌约简,揭示其特性及精度与效率权衡见解,为相关研究奠定基础。

中文摘要 AI 辅助

近期基于视觉Transformer的视觉场所识别(VPR)方法,尤其是基础模型,取得了显著识别性能。但这些模型在整个网络中处理所有视觉令牌,导致大量计算开销,阻碍其在实时和资源受限场景中的部署。为此提出首个用于高效视觉场所识别的令牌约简系统基准。该基准在多个先进VPR模型及涵盖城市、郊区和自然环境的多样基准数据集上,全面评估代表性令牌剪枝、合并及混合剪枝-合并方法。还从不同约简配置下的识别性能、计算复杂度、推理速度、定性可视化及边缘设备部署效率等多视角研究令牌约简。通过广泛实验和深入分析,揭示了VPR中令牌约简的多个重要特性,并提供了精度与推理效率权衡的实用见解。例如,令牌约简可将计算成本降低多达29%,吞吐量提高多达44%,同时识别准确率下降不到1%。总体而言,这项工作为未来令牌高效VPR和高效视觉检索系统研究奠定了全面基础。代码和模型将在指定网址提供。

英文摘要

Recent visual place recognition (VPR) methods based on vision transformers, particularly foundation models, have achieved remarkable recognition performance. However, these models process all visual tokens throughout the entire network, resulting in substantial computational overhead, which hinders their deployment in real-time and resource-constrained scenarios. A natural question thus arises: are all visual tokens necessary for VPR? To answer this question, we present the first systematic benchmark of token reduction for efficient visual place recognition. Our benchmark comprehensively evaluates representative token pruning, token merging, and hybrid pruning-merging methods across multiple state-of-the-art VPR models and diverse benchmark datasets covering urban, suburban, and natural environments. We further investigate token reduction from multiple perspectives, including recognition performance under different reduction configurations, computational complexity, inference speed, qualitative visualization, and deployment efficiency on edge devices. Through extensive experiments and in-depth analysis, our benchmark reveals multiple important characteristics of token reduction in VPR and provides several practical insights into the trade-offs between accuracy and inference efficiency. For example, token reduction can reduce computational cost by up to 29\% and improve throughput by up to 44\%, while incurring less than 1\% degradation in recognition accuracy. Overall, this work establishes a comprehensive foundation for future research on token-efficient VPR and efficient visual retrieval systems. Our codes and models will be available at https://github.com/Tong-Jin01/TokenReduction4VPR

发表机构

  • Shenyang Institute of Automation, Chinese Academy of Sciences(中国科学院沈阳自动化研究所)
  • University of Chinese Academy of Sciences(中国科学院大学)
  • Shenzhen University of Advanced Technology(深圳先进技术大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑