发表机构
The Hebrew University of Jerusalem(耶路撒冷希伯来大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对冷启动半监督学习中分类器自监督不适定的问题,提出VAST方法,通过解耦伪标签生成与训练,利用冻结嵌入几何推断信念并蒸馏至归纳分类器,在多个数据集上显著优于现有基线。
AI 中文摘要
现代半监督学习(SSL)将伪标签生成与分类器训练耦合在一起,利用分类器自身的置信度来选择伪标签,进而用于更新模型。在冷启动场景下(每个类别最多只有少量标签可用),这种耦合是不适定的,因为分类器在尚未学习之前无法自我监督。为解决此问题,我们提出了VAST(真实性感知的半监督训练),该方法将这两个阶段解耦。首先直接从冻结的自监督嵌入的几何结构中推断出未标记集上的概率信念,然后仅将其蒸馏到归纳式分类器中。该构造基于真实性矩阵(Veracity Matrix),这是一种基于核的结构,它聚合了数据流形上的标签证据,并可在逐观测加权似然模型下解释为狄利克雷后验。此外,我们引入了真实性传播(Veracity Propagation),这是一种自终止的信念传播步骤,可将覆盖范围扩展到标记集的核邻域之外。在一个受控协议下(所有方法接收相同的冻结嵌入和标记集),VAST在三个数据集上的每个操作点均优于最强的基于图的SSL基线,在9项比较中有7项具有统计学显著提升,同时产生了一个可部署的归纳式分类器,而非需要转导式图推断。与端到端的置信门控SSL相比,我们进一步发现这些方法在此设置下表现不佳,并且在我们的实验中,它们并未持续超过仅使用标记数据的性能。
英文摘要
Modern semi-supervised learning (SSL) couples pseudo-label generation and classifier training, using the classifier's own confidence to select the pseudo-labels that are then used to update the model. In the cold-start regime, where at most a few labels per class are available, this coupling is ill-posed, since the classifier cannot supervise itself before it has learned. To address this problem, we propose VAST (Veracity-Aware Semi-Supervised Training), which decouples these two stages. Probabilistic beliefs over the unlabeled set are first inferred directly from the geometry of a frozen self-supervised embedding and only then distilled into an inductive classifier. The construction rests on the Veracity Matrix, a kernel-based structure that aggregates label evidence across the data manifold and admits an interpretation as a Dirichlet posterior under a per-observation powered-likelihood model. Additionally, we introduce Veracity Propagation, a self-terminating belief-spreading step that extends coverage beyond the kernel neighborhood of the labeled set. Under a controlled protocol in which all methods receive identical frozen embeddings and labeled sets, VAST outperforms the strongest graph-based SSL baselines at every operating point across three datasets, with statistically significant gains in 7 of 9 comparisons, while producing a deployable inductive classifier rather than requiring transductive graph inference. Compared with end-to-end confidence-gated SSL, we further find that these methods underperform in this setting and, in our experiments, do not consistently exceed labeled-only performance.