发表机构
NTT, Inc.(日本电信电话株式会社)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对UU学习在分布偏移下的挑战,提出基于重要性加权的适应方法,利用训练与少量测试UU数据,统一处理PU和噪声标签学习,无需偏移类型假设,实验验证有效。
AI 中文摘要
无标注-无标注(UU)学习允许我们从两个具有不同类别先验的无标注数据集中学习一个二元分类器。它是一个通用框架,因为它涵盖了多种监督学习,如正例-无标注(PU)学习、噪声标签学习和基于相似性的学习。现有的UU学习假设测试和训练分布具有相同的类条件密度。然而,由于分布偏移,这一假设在实践中很少成立。本文提出了一种针对UU学习的分布偏移适应方法,该方法使用训练分布中的UU数据和测试分布中的少量UU数据。所提出的方法基于重要性加权,通过使用带有估计重要性权重的训练数据来最小化测试风险。尽管现有的重要性加权方法无法处理UU数据,但我们展示了可以以原则性的方式实现这一点。得益于UU学习的通用性,我们的方法可以在一个统一框架内处理各种学习问题,如分布偏移下的PU学习和噪声标签学习,而现有方法通常针对特定问题进行定制。此外,它不需要对偏移类型(如协变量偏移)做任何假设。我们通过真实世界数据集实验证明了所提出方法的有效性。
英文摘要
Unlabeled-unlabeled (UU) learning allows us to learn a binary classifier from two sets of unlabeled data with different class-priors. It is a general framework because it includes a wide variety of supervised learning such as positive-unlabeled (PU) learning, noisy label learning, and similarity-based learning. Existing UU learning assumes that the test and training distributions have the same class-conditional densities. However, this assumption rarely holds in practice due to distribution shifts. This paper proposes a distribution shift adaptation method for UU learning that uses UU data in the training distribution and a few UU data in the test distribution. The proposed method is based on the importance weighting, which minimizes the test risk by using training data with estimated importance weights. Although existing importance weighting methods cannot handle UU data, we show that it can be done in a principled manner. Thanks to the generality of UU learning, our method can handle various learning problems such as PU and noisy label learning under distribution shift within a single framework while existing methods are usually tailored to a specific problem. Moreover, it does not require any assumption of the shift types such as covariate shift. We experimentally demonstrate the effectiveness of the proposed method with real-world datasets.
Comments19 pages