发表机构
Università degli Studi di Milano; Istituto Nazionale di Fisica Nucleare(米兰大学; 意大利国家核物理研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究从多个无标签混合物进行多类学习的问题,利用后验单纯形几何结构,提出无先验程序,通过训练标准分类器区分混合物身份并提取潜在类结构,实验证明能恢复潜在类及其比例,缩小弱监督与全监督性能差距。
AI 中文摘要
在许多分类问题中,可靠的实例级标签不可用。但可构建弱富集无标签样本。无标签分类(CWoLa)表明在二分类中,训练区分不同类比例的不纯混合物的分类器可恢复最优类判别器。本文将此原理扩展到多类学习,从多个无标签混合物中学习。证明了贝叶斯最优混合物分类器将数据点映射到混合物后验空间中的\((K - 1)\) - 单纯形。利用此几何结构,提出无先验程序,通过后验单纯形拟合或瓶颈架构训练标准分类器区分混合物身份并提取潜在类结构。在MNIST、CIFAR - 10和Galaxy10 DECaLS上的实验表明,仅混合物身份就能恢复潜在类及其在混合物中的比例。缩小了弱监督和全监督性能之间的差距,为标签稀缺领域的多类发现提供了数学基础且可扩展的工具。
英文摘要
In many classification problems, reliable instance-level labels are unavailable. However, it is often possible to construct weakly enriched unlabeled samples: datasets selected by different cuts, sources, populations, or experimental conditions that change latent class proportions without revealing them. Classification without Labels (CWoLa) shows that, in the binary case ($K=2$), a classifier trained to distinguish two impure mixtures with different class proportions can recover an optimal class discriminator without knowing the mixture proportions. We extend this principle to multiclass learning from several unlabeled mixtures ($K>2$), where the learner observes only mixture identity and neither latent class labels nor class-prior matrices. We prove that, for a multiclass mixture model, the Bayes-optimal mixture classifier $g^\star$ maps data points into a $(K-1)$-simplex embedded in mixture-posterior space. The $K$ vertices of this simplex are induced by the latent classes through the unknown mixing matrix. Leveraging this geometry, we propose prior-free procedures that train a standard classifier to distinguish mixture identities and then extract latent class structure using either post-hoc simplex fitting or a bottleneck architecture. Experiments on MNIST, CIFAR-10, and Galaxy10 DECaLS show that mixture identity alone can recover latent classes and their fractions in the mixture. By narrowing the gap between weakly supervised and fully supervised performance, we provide a mathematically grounded, scalable tool for multiclass discovery in label-scarce domains.