聚合邻域嵌入投影与基于排序的流形学习用于图像检索
Aggregating Neighbor Embedding Projection and Rank-Based Manifold Learning for Image Retrieval
浏览论文内容
中文总结 AI 辅助
该研究针对图像检索中高维空间排序难的问题,提出结合UMAP投影与Borda计数聚合的邻域嵌入投影及排序流形学习框架,在多数据集上提升了检索精度与MAP值。
中文摘要 AI 辅助
基于内容的图像检索(CBIR)在深度学习的推动下已取得显著进展,但在高维特征空间中对相似图像进行有效排序仍具挑战性,此时成对距离往往无法捕捉上下文关系,且视觉特征与高层概念之间的语义鸿沟依然存在。流形学习与基于排序的优化方法已成为互补策略,分别用于改进特征表示和利用排序列表中嵌入的上下文信息(如图像间的邻域关系)。然而,将这些基于投影的策略与基于排序的策略相结合以利用其互补特性仍是一个具有挑战性的研究问题。为解决该问题,我们提出了一种通过排序聚合将邻域嵌入投影与基于排序的流形学习相结合的框架。Uniform Manifold Approximation and Projection(UMAP,均匀流形近似与投影)生成替代的低维特征表示,将从UMAP投影得到的排序列表与基于排序的重排序方法得到的排序列表通过Borda Count(博达计数)聚合策略进行组合。实验在多个公开数据集上开展,使用从ResNet152、Swin Transformer和DINOv2模型中提取的深度学习特征。结果表明,所提方法在多种场景下提升了检索有效性,尤其是当基线表示难以达到高准确率时。该聚合策略还常能提升排名靠前位置的质量,在不同数据集和特征提取器上均获得了具有竞争力的平均精度均值(MAP)和精度值。这些发现表明,通过排序聚合将基于投影的流形学习策略与基于排序的流形学习策略相结合,可为图像检索任务提供互补的上下文信息。
英文摘要
Content-based image retrieval (CBIR) has advanced significantly with deep learning, yet effectively ranking similar images remains challenging, particularly in high-dimensional feature spaces, where pairwise distances often fail to capture contextual relationships and the semantic gap between visual features and high-level concepts persists. Manifold learning and rank-based refinement methods have emerged as complementary strategies, respectively improving feature representations and exploiting contextual information embedded in ranked lists, such as neighborhood relationships among images. However, combining these projection-based and rank-based strategies to exploit their complementary properties remains a challenging research problem. To address this, we propose a framework that combines neighbor embedding projections with rank-based manifold learning through rank aggregation. Uniform Manifold Approximation and Projection (UMAP) generates alternative low-dimensional feature representations, and ranked lists obtained from UMAP projections and rank-based re-ranking methods are combined using the Borda Count aggregation strategy. Experiments were conducted on several public datasets using deep learning features extracted from ResNet152, Swin Transformer, and DINOv2 models. Results show that the proposed approach improves retrieval effectiveness in several scenarios, particularly when the baseline representation struggles to achieve high precision. The aggregation strategy also often improves the quality of top-ranked positions, leading to competitive Mean Average Precision (MAP) and Precision values across different datasets and feature extractors. These findings suggest that combining projection-based and rank-based manifold learning strategies through rank aggregation can provide complementary contextual information for image retrieval tasks.
发表机构
- São Paulo State University (UNESP)(圣保罗州立大学)
- University of São Paulo (USP)(圣保罗大学)
机构由 AI 辅助整理,请以论文原文为准。