arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在概率对比学习中,共享温度是否意味着共享角度尺度?

Does a Shared Temperature Imply a Shared Angular Scale in Probabilistic Contrastive Learning?

Ningkang Peng, Qianfeng Yu, Jingyang Mao, Xiaoqian Peng, Tingyu Lu, Peirong Ma, Yanhui Gu

arXiv 2609.38784首次发表:更新:

发表机构

Nanjing Normal University; Nanjing University of Chinese Medicine; Tohoku University(南京师范大学; 南京中医药大学; 东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究证明在概率对比学习中,共享温度并不直接对应共享角度尺度,并提出了vMF得分的前导角度增益理论,通过实验验证了其在多种数据集上的预测准确性,并展示了纯角度编辑对表示学习的改善作用。

AI 中文摘要

在概率对比学习中,共享温度通常被解释为共享相似度尺度,但这种解释对于高维分布类表示并不成立。我们研究了当表示维度和类浓度共同增长时,ProCo所使用的精确von Mises-Fisher(vMF)概率得分。我们证明了该得分保留了一个类依赖的前导角度增益$g_c=A_c/\ au$,其中$A_c$是平均合成长度。该增益进入Softmax竞争、成对决策边界和特征梯度。在真实的CIFAR-LT、ImageNet-LT和iNaturalist表示上,该理论准确预测了完整vMF得分下的边界移动和局部梯度变化。类级温度调整也改变了余弦零截距和有限维响应。我们构建了保持截距和纯角度控制,以将前导增益与这些伴随变化分开。完全增益均衡在首阶产生共享尺度余弦原型规则;有限维边际条件保证两个分类器的一致性。在16个冻结表示设置中,预测一致性为98.43%-99.99%,不一致集中在小的余弦边际处。在受控的仅对比训练中,使用训练频率先验,纯角度编辑在所有测试的CIFAR-10/100不平衡因子下改善了两种学习表示,并在ImageNet-LT上保持积极变化。因此,vMF浓度不仅描述了类分布,还在高维概率对比学习中形成了决策和学习尺度。

英文摘要

In probabilistic contrastive learning, a shared temperature is commonly interpreted as a shared similarity scale, but this interpretation does not hold for high-dimensional distributional class representations. We study the exact von Mises-Fisher (vMF) probabilistic score used by ProCo when representation dimension and class concentration grow jointly. We prove that the score retains a class-dependent leading angular gain $g_c=A_c/τ$, where $A_c$ is the mean resultant length. This gain enters Softmax competition, pairwise decision boundaries, and feature gradients. On real CIFAR-LT, ImageNet-LT, and iNaturalist representations, the theory accurately predicts boundary movements and local gradient changes under the full vMF score. Classwise temperature adjustment also changes the cosine-zero intercept and finite-dimensional response. We construct intercept-preserving and Pure Angular controls to separate the leading gain from these accompanying changes. Complete gain equalization yields a shared-scale cosine prototype rule at leading order; a finite-dimensional margin condition guarantees agreement of the two classifiers. Across 16 frozen representation settings, prediction agreement is 98.43-99.99%, with disagreements concentrated at small cosine margins. In controlled contrastive-only training with the training-frequency prior, Pure Angular editing improves both learned representations at all tested CIFAR-10/100 imbalance factors and retains positive changes on ImageNet-LT. Thus vMF concentration not only describes class distributions, but also forms a decision and learning scale in high-dimensional probabilistic contrastive learning.

Comments59 pages, including supplementary material

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑