AI 中文总结
本文提出基于嵌套再生核希尔伯特空间的统一学习框架,通过密度估计与贝叶斯分类应用,实现监督与无监督学习的强收敛保证。
AI 中文摘要
我们引入了一个基于再生核希尔伯特空间(RKHSs)的统一框架,用于监督与无监督学习。我们提出了嵌套RKHS的概念,即一个递增的RKHS序列,其并集在环境$L^2$空间中稠密。这一构造使得在每一个有限逼近层次上利用RKHS方法的计算结构成为可能,同时保留完整$L^2$空间的逼近能力。我们首先研究固定RKHS中的统计学习。在无监督和监督两种设置下,我们建立了在连续嵌入$L^2$的RKHS中取值的随机元素的经验平均值的强收敛结果。然后,我们利用嵌套RKHS结构,将这些统计收敛结果与RKHS并集在环境$L^2$空间中的稠密性相结合。我们将此构造应用于概率密度估计和二元贝叶斯分类。对于密度估计,我们构造了目标密度在RKHS上的正交投影的显式经验估计量,并建立了它们的几乎必然收敛性。随后,一个对角选择程序产生一个估计量序列,该序列在$L^2$中几乎必然收敛于目标密度。对于二元贝叶斯分类,我们将贝叶斯目标的逼近表述为适当的总体和经验风险泛函的最小化。我们证明了经验最小化器的一致性,并利用嵌套结构获得在$L^2$中向贝叶斯目标的收敛。我们进一步提供了在环境$L^2$空间中关于超平面和加权距离的几何解释。最后,我们表明总体风险在误分类概率方面具有直接的概率解释。
英文摘要
We introduce a unified framework for supervised and unsupervised learning through density estimation using reproducing kernel Hilbert spaces (RKHSs). To this end, we introduce the notion of a nested RKHS, that is, an increasing sequence of RKHSs whose union is dense in an ambient $L^2$-space. This construction allows us to exploit the computational structure of RKHS methods at each finite approximation level while preserving the approximation power of the entire $L^2$-space. We first study statistical learning in a fixed RKHS. In both unsupervised and supervised settings, we establish strong convergence results for empirical averages of random elements taking values in an RKHS continuously embedded in $L^2$. We then use the nested RKHS structure to combine these statistical convergence results with the density of the union of the RKHSs in the ambient $L^2$-space. We apply this construction to probability density estimation and binary Bayesian classification. For density estimation, we construct explicit empirical estimators of the orthogonal projections of the target density onto the RKHSs and establish their almost sure convergence. A diagonal selection procedure then yields a sequence of estimators converging almost surely to the target density in $L^2$. For binary Bayesian classification, we formulate the approximation of the Bayes target as the minimization of suitable population and empirical risk functionals. We prove consistency of the empirical minimizers and use the nested structure to obtain convergence towards the Bayes target in $L^2$. We further provide a geometric interpretation in terms of hyperplanes and weighted distances in the ambient $L^2$-space. Finally, we show that the population risk admits a direct probabilistic interpretation in terms of the probability of misclassification.
CommentsMinor revisions: corrected typos and clarified a definition and a statement