发表机构
Czech Technical University in Prague(布拉格捷克理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出全局平均精度(gAP)的可微代理gSAP,通过全局排序所有查询-候选对来改进表示学习,在监督度量学习、跨模态对齐和自监督预训练中优于现有损失,并提升检索与分类性能。
AI 中文摘要
标准信息检索指标(如平均精度均值(mAP))一次评估一个查询的性能,基于查询与其正样本之间的相似度与负样本之间的相似度进行比较。常见的表示学习损失(如InfoNCE和每查询AP代理)也是如此。它们都没有考虑相似度是否跨查询可比,而任何依赖单一决策阈值的系统都需要这种可比性。全局平均精度(gAP)通过将所有查询-候选对排列在一个列表中并计算单个AP来实现这一点。我们引入了gSAP,这是gAP的可微代理。它只需要一个相似度矩阵和一个标记正对的二元矩阵,与现有损失相同的输入,因此它是它们的即插即用替代品,并且对编码器、模态和监督来源不可知。由于它同时考虑批次中所有可能的成对比较,它即使在低温下也能保持可训练,而每查询代理在这种状态下会耗尽梯度。将其替换到既定配方中可改善监督度量学习、跨模态对齐和自监督预训练,据我们所知,它是第一个在后两者中取代社区标准InfoNCE的排序损失。其相似度在查询间更一致,这推动了在通用阈值下的性能提升。gSAP在相同精度下检索到的正对数量是最强AP代理的四倍,并且在数据库中添加无正样本的查询时退化最少。除了阈值化之外,使用gSAP训练的模型还能学习到更好的表示,具有更高的迁移、$k$NN和零样本分类准确率。
英文摘要
Standard information retrieval metrics, such as mean Average Precision (mAP), assess performance one query at a time, based on how the similarities between a query and its positives compare against those with its negatives. The same holds for common representation learning losses, such as InfoNCE and per-query AP surrogates. None of them considers whether similarities are comparable across queries, which any system with a single decision threshold relies on. Global Average Precision (gAP) does, by ranking all query-candidate pairs in one list and computing a single AP. We introduce gSAP, a differentiable surrogate of gAP. It needs only a similarity matrix and a binary matrix marking the positive pairs, the same input as existing losses, so it is a drop-in replacement for them and agnostic to the encoder, the modality, and the source of supervision. Since it considers all possible pairwise comparisons in the batch jointly, it also remains trainable at low temperatures, a regime where per-query surrogates run out of gradient. Swapping it into established recipes improves supervised metric learning, cross-modal alignment, and self-supervised pretraining, where, to our knowledge, it is the first ranking loss to replace the community standard InfoNCE in the latter two. Its similarities are more consistent across queries, which drives the gains under a universal threshold. gSAP retrieves up to four times as many positive pairs as the strongest AP surrogate at the same precision, and it degrades the least when queries with no positives in the database are added. Beyond thresholding, models trained with gSAP also learn better representations, with higher transfer, $k$NN and zero-shot classification accuracy.