发表机构
University of Kentucky; Kentucky Geological Survey(肯塔基大学; 肯塔基地质调查局)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究自监督视觉表征学习中预训练-微调与联合训练两种范式,通过在多种任务及不同标记数据比例下评估性能,发现JT在低标记时提高效率且稳健,PFT在专业领域更可靠,为混合自监督半监督学习提供基准和指导。
AI 中文摘要
自监督是从未标记数据中学习视觉表征的强大技术。现有技术主要采用两阶段自监督学习方法,先在未标记数据上预训练,再在标记数据上微调。虽已证明其有效性,但自监督与监督学习目标间的交互仍未被充分理解。本文系统研究训练期间联合优化这两个目标是否更好。比较了预训练-微调(PFT)和联合训练(JT)两种范式,在多种计算机视觉任务及不同比例标记数据下评估性能。结果表明,PFT和JT的相对有效性强烈依赖于任务、标记数据可用性和领域复杂性。JT在低标记设置下提高数据和训练效率且稳健,PFT在更专业领域更可靠。还分析了表征质量等,为基于混合自监督的半监督学习建立了综合实证基准并提供实用指导。
英文摘要
Self-supervision is a powerful technique for learning visual representations from unlabeled data. Existing techniques primarily adopt a two-stage approach for self-supervised learning (SSL): a pretraining stage on unlabeled data followed by a finetuning stage on labeled data. While this pipeline has demonstrated extreme effectiveness, the interaction between self-supervised and supervised learning objectives remains insufficiently understood. In this work, we systematically investigate whether jointly optimizing the self-supervised and supervised objectives during training provides a better alternative. We compare two training paradigms: (1) the aforementioned pretraining followed by finetuning (PFT) and (2) joint training (JT), where self-supervised and supervised losses are optimized simultaneously in the same network. Across eight representative SSL methods and diverse computer vision tasks on natural, medical, crisis response, and remote sensing data, we evaluate performance under varying percentages of labeled data. Our results reveal that the relative effectiveness of PFT and JT depends strongly on the task at hand, the availability of labeled data, and the complexity of the domain. We find that JT consistently improves data and training efficiency while being robust in low-label settings, while PFT is more reliable in more specialized domains. We further analyze representation quality, robustness, and cross-domain generalization, providing new insights into how self-supervised and supervised objectives interact during optimization. We establish a comprehensive empirical benchmark for hybrid SSL-based semi-supervised learning and offer practical guidance for selecting appropriate training strategies across diverse vision applications.