AI 中文总结
针对多个特征空间异构但标签相同的联盟数据集,提出合并特征空间并利用矩阵补全构建统一数据集的方法,实验证明统一训练的分类器性能优于单独训练。
AI 中文摘要
在许多应用领域,如学生辍学、保险欺诈、贷款审批和机器故障中,存在多个带标签的公开数据集,这些数据集具有以下特点:(i) 数据涉及相同类型的对象,但实际底层对象集合互不相交;(ii) 类别标签相同;(iii) 数据集的特征空间大部分不同(异构),仅有少数共享特征。我们将此类数据集称为联盟数据集。单个分类器无法同时在两个数据集上训练,而一个数据集上训练的分类器也无法在另一个数据集上测试。在本文中,我们提出一种方法,将一对给定的异构联盟数据集的特征空间合并为单一特征空间。然后,我们使用矩阵补全方法,基于合并后的特征空间创建统一数据集。假设是,合并后的表示有助于分类知识从一个数据集迁移到另一个数据集。我们在多对联盟异构数据集上使用多种分类器进行实验,结果表明,在多个联盟数据集对上,任何在统一表示上训练的分类器总是优于分别在组成联盟数据集上训练的分类器。这项工作提供了一种简单的方法,通过统一并同时使用多个联盟数据集来显著提高分类器性能。
英文摘要
In many application domains, such as student dropout, insurance fraud, loan approval, and machine failures, several labelled public datasets are available where (i) data is about the same type of objects but the set of actual underlying objects are disjoint; and (ii) the class labels are same; and (iii) the feature spaces of the datasets are largely distinct (heterogeneous), with a few shared features. We call such datasets as allied. A single classifier cannot be trained on both datasets together, and one classifier trained on one dataset cannot be tested on the other. In this paper, we propose a method to merge the feature-spaces into a single feature-space for a pair of given allied heterogeneous datasets. We then use a matrix completion method to create a unified dataset based on the merged feature-space. The hypothesis is that the merged representation facilitates the transfer of classification knowledge from one dataset to another. We conduct experiments on several pairs of allied, heterogeneous datasets and several classifiers to demonstrate that any classifier trained on the unified representation always outperforms classifiers separately trained on the constituent allied datasets on several pairs of allied datasets. This work provides an easy way to substantially improve classifier performance by unifying and using multiple allied datasets together.