发表机构
Budapest University of Technology and Economics(布达佩斯技术与经济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文扩展一般化朴素贝叶斯至连续特征,证明其结构学习仅依赖成对copula,并基于拟阵提出贪心算法以最小化KL散度,涵盖高斯、高斯copula及任意分布三种情况,最后通过GNB森林简化模型并在真实数据上验证性能。
AI 中文摘要
一般化朴素贝叶斯(GNB)模型最初是针对离散和分类随机变量提出的,作为经典朴素贝叶斯的扩展。我们现在将GNB框架适配到连续解释变量。本文的一个核心结果是,GNB的结构学习仅依赖于二元边际的成对copula。我们证明了GNB结构可以赋予拟阵的基,因此我们给出了在训练数据上寻找最优GNB结构的贪心算法,其意义在于最小化Kullback-Leibler散度。考虑了三种情况:联合高斯分布;然后是一个更灵活的模型,其中我们假设依赖结构由具有任意边际的高斯copula描述;以及一个更灵活的情况,即联合连续概率分布是任意的,即copula和边际分布都是任意的。基于新引入的GNB森林概念,给出了一种模型简化方法。最后,我们将新引入的GNB分类结果与真实数据集上的其他经典“玻璃箱”算法进行了比较。
英文摘要
The Generalized Naive Bayes (GNB) model was introduced for discrete and categorical random variables as an extension of classic Naive Bayes. We now accommodate the GNB framework to continuous explanatory variables. A central result of the paper is that structure learning of the GNB depends only on the pair copulas of the bi-variate marginals. We proved that the GNB structure can be assigned to the basis of a matroid, therefore we give greedy algorithms for finding the optimal GNB structure on the training data, in sense of minimizing Kullback-Leibler divergence. Three cases are considered: joint Gaussian distribution, then a more flexible model where we suppose the dependence structure to be described by a Gaussian copula with arbitrary marginals, and an even more flexible case where the joint continuous probability distribution is arbitrary, i.e. copula and marginal distributions are arbitrary. A method for model reduction, based on the newly introduced concept of GNB forest is given. We close the paper by comparing the newly introduced GNB classification results to other classical "glass-box" algorithms on real datasets.