发表机构
CNRS; Inria; Institut Universitaire de France; Univ Rennes(法国国家科学研究中心; 法国国家信息与自动化研究所; 法国大学研究院; 雷恩大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出去随机化框架,利用分解PAC-贝叶斯理论将随机多数投票的保证转化为确定性多数投票的证书,推导两类泛化界并引出自界学习算法。
AI 中文摘要
加权多数投票是许多成功集成方法的核心。PAC-贝叶斯理论通过分析随机分类器的期望风险,为这类模型提供了紧密的泛化保证,而分析确定性多数投票的风险则依赖于替代界。为避免这些替代,Zantedeschi等人(2021)引入了针对随机多数投票的保证,但由此产生的模型仍是随机的。在本文中,我们提出了一个针对随机多数投票的去随机化框架。为此,我们将分解PAC-贝叶斯理论的最新进展直接应用于多数投票权重向量空间,将随机保证转化为单个确定性多数投票的证书。我们推导了两族高概率泛化界,涵盖集成构造的数据无关与数据相关情形,这自然引出了一个优化确定性多数投票保证的自界学习算法。
英文摘要
Weighted majority votes are central to many successful ensemble methods. PAC-Bayesian theory provides tight generalization guarantees for such models by analyzing the expected risk of stochastic classifiers, while analyzing the risk of deterministic majority votes relies on surrogate bounds. To avoid these surrogates, Zantedeschi et al. ( 2021) introduced guarantees for stochastic majority votes, but the resulting models remain randomized. In this paper, we propose a derandomization framework for stochastic majority votes. To do so, we apply recent advances in disintegrated PAC-Bayesian theory directly to the space of majority vote weight vectors, transforming stochastic guarantees into certificates for a single deterministic majority vote. We derive two families of high-probability generalization bounds, covering both data-independent and data-dependent constructions of the ensemble, which naturally lead to a self-bounding learning algorithm optimizing deterministic majority vote guarantees.