发表机构
University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究评估贪心K中心在多种度量空间中用于主动学习选择的性能,发现将未标注实例映射到带熵加权的预测概率空间时效果最优。
AI 中文摘要
随着近年来的快速发展,许多不同的方法现已广泛应用于分类任务。然而,训练这些模型需要大量标注数据,主动学习是解决该问题的潜在方案。基于池的主动学习仅从未标注数据集中查询最具信息性的样本以最小化成本,而基于多样性的方法则旨在选择数据的代表性子集。确定选择过程的目标有多种,包括精确K中心、精确K中位数和贪心K中心。本文中,我们将重点评估贪心K中心在多种度量空间中的性能:原始特征空间、线性判别分析(LDA)空间以及模型导出的概率空间(含和不含基于熵的加权)。以随机森林分类器作为基线评估器,我们在合成数据集和真实世界数据集上的实证结果表明,将未标注实例映射到预测概率空间并通过熵对结果加权,在使用贪心K中心进行主动学习选择时,通常优于其他选项。
英文摘要
With rapid advancement over the last few years, many different methods are now widely used for classification. However, training these models requires substantial labeled data. Active Learning is a potential solution to this problem. Pool-based active learning minimizes costs by querying only the most informative samples from an unlabeled dataset. Diversity-based approaches, on the other hand, attempt to select a representative subset of the data. There are many different objectives for determining the selection process, including exact K-center, exact K-median, and Greedy K-center. In this paper, we will focus on evaluating the performance of Greedy K-center across a variety of metric spaces: the raw feature space, a Linear Discriminant Analysis (LDA) space, and a model-derived probability space (with and without entropy-based weighting). Using Random Forest classifiers as a baseline evaluator, our empirical results on synthetic and real-world datasets demonstrate that mapping unlabeled instances into a predictive probability space and weighting the result by entropy often dominates the other options for active learning selection with Greedy K-center.