arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语义增益从何而来?面向长尾识别的语义知识驱动对比学习的复现与扩展

Where Does the Semantic Gain Come From? A Reproduction and Extension of Semantic Knowledge-driven Contrastive Learning for Long-Tailed Recognition

Sushrut Ghimire

arXiv 2610.04104首次发表:更新:

发表机构

Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU)(弗里德里希-亚历山大大学埃尔朗根-纽伦堡)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究复现并扩展了SKCL,发现其声称的增益主要来自ConCutMix和更长训练,而非语义图,并提出了结合CutMix分支等改进。

AI 中文摘要

语义知识驱动对比学习(SKCL)使用语言模型判断哪些类别相关,并将每张图像拉向其语义邻居的原型。在CIFAR-100-LT(beta=100)上,它报告了54.02%的top-1准确率,比其构建基础的平衡对比学习(BCL)高出2.01个百分点。代码和类别描述未公开。我在一个框架中重新实现了SKCL、BCL和ConCutMix,对照公开基线代码进行验证,并使用三个随机种子运行每种配置。两个基线在1.5个百分点内复现,但按论文描述构建于BCL之上的SKCL最终比BCL低0.56个百分点。为查明原因,我将SKCL添加到作者自己的ConCutMix代码中。按论文的300轮训练,它达到53.69,仅比已发表数字低0.33。然而,在同一预算下,语义图仅比ConCutMix增加0.23个百分点,而将ConCutMix多训练100轮则增加1.04个百分点。结合ConCutMix已发表的领先BCL的1.15个百分点,这解释了声称的增益。一个从未见过该图的BCL模型,其最混淆的类别中已有41.2%与图的top-2邻居重合(随机概率为2.0%),这说明了为何该图在这些基准上增益甚微。我还测试了SKCL的若干变体。将其与CutMix分支结合可提升1.06个百分点,图的自适应版本略有改进(在两个代码库中分别提升+0.34和+0.28),尽管这些增益在随机种子噪声范围内。

英文摘要

Semantic Knowledge-driven Contrastive Learning (SKCL) uses a language model to decide which classes are related, and pulls each image towards the prototypes of its semantic neighbours. On CIFAR-100-LT (beta = 100) it reports 54.02% top-1 accuracy, 2.01 points above Balanced Contrastive Learning (BCL), the method it builds on. The code and the class descriptions are not public. I reimplement SKCL, BCL and ConCutMix in one framework, check it against the public baseline code, and run every configuration with three seeds. The two baselines reproduce within 1.5 points, but SKCL built on BCL, as the paper describes it, ends up 0.56 points below BCL. To find out why, I add SKCL to the authors' own ConCutMix code. Trained for the paper's 300 epochs, it reaches 53.69, only 0.33 below the published number. At the same budget, however, the semantic graph adds just 0.23 points over ConCutMix, while training ConCutMix for 100 more epochs adds 1.04. Together with ConCutMix's published lead over BCL (1.15), this explains the claimed gain. A BCL model that never sees the graph already shares 41.2% of the graph's top-2 neighbours with its own most-confused classes (2.0% by chance), which shows why the graph adds so little on these benchmarks. I also test several changes to SKCL. Combining it with the CutMix branch improves it by 1.06 points, and an adaptive version of the graph improves it slightly (+0.34 and +0.28 in two codebases), although these gains are within seed noise.

Comments11 pages, 4 figures. Code: https://github.com/sushrutghimire1/Reproduction-and-extension-of-SKCL

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑