发表机构
Van Lang University; Ho Chi Minh City University of Transport(范朗大学; 胡志明市交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出HEB-NB及HEB-AODE,解决NB固定平滑强度致高基数表格数据偏差问题,经31个基准测试,其在概率指标、对数损失及ECE上均有显著改进。
AI 中文摘要
朴素贝叶斯(NB)分类器仍是分类数据的标准选择,但其广泛使用的平滑规则(如拉普拉斯、利德斯通、克里切夫斯基-特罗菲莫夫和m估计)均规定固定平滑强度,忽略特征基数、样本量和类别不平衡,在现代高基数表格数据上引入非零偏差。我们提出分层经验贝叶斯朴素贝叶斯(HEB-NB),其中每个类别-特征条件概率由狄利克雷先验平滑,该先验的浓度通过第二类最大似然数据自适应学习,实现跨类别的合理信息共享,同时保留闭式推理。我们进一步引入HEB平均单依赖估计器(HEB-AODE),表明自适应平滑可清晰迁移至NB的结构松弛。理论上,我们为HEB-NB建立非渐近ℓ₁误差界,匹配经验分布极小极大率加非零数据自适应偏差,同时得到匹配拉普拉斯的紧下界,产生与拉普拉斯的有限样本风险级严格分离。我们还通过总变差张量化推导插件式 excess贝叶斯风险界,并得到总体Top-1预期校准误差(ECE)推论。实证上,在31个UCI和OpenML基准中,HEB-NB在概率指标上取得最佳平均弗里德曼秩,在高基数数据集上对数损失降低达22.1%,且HEB-AODE较普通AODE持续改进。结合HEB-NB与互信息加权使Top-1 ECE降低41%-70%,在概率准确性和校准上取得显著提升。
英文摘要
The Naive Bayes (NB) classifier remains a standard choice for categorical data, yet its widely used smoothing rules, such as Laplace, Lidstone, Krichevsky-Trofimov, and the $m$-estimate, all prescribe a fixed smoothing strength that ignores feature cardinality, sample size, and class imbalance, inducing a non-vanishing bias on modern high-cardinality tabular data. We propose hierarchical empirical-Bayes Naive Bayes (HEB-NB), in which each class-feature conditional probability is smoothed by a Dirichlet prior whose concentration is learned data-adaptively via Type-II maximum likelihood, enabling principled information sharing across classes while retaining closed-form inference. We further introduce HEB average one-dependence estimators (HEB-AODE), showing that the adaptive smoothing transfers cleanly to structural relaxations of NB. Theoretically, we establish a non-asymptotic $\ell_1$ error bound for HEB-NB matching the empirical-distribution minimax rate plus a vanishing data-adaptive bias, together with a matching Laplace-tight lower bound that yields a finite-sample, risk-level strict separation from Laplace. We further derive a plug-in excess Bayes-risk bound via total-variation tensorization and a population top-1 expected calibration error (ECE) corollary. Empirically, across 31 UCI and OpenML benchmarks, HEB-NB attains the best average Friedman rank on probabilistic metrics, with up to 22.1% log-loss reductions on high-cardinality datasets and consistent improvements of HEB-AODE over vanilla AODE. Combining HEB-NB with mutual-information weighting reduces top-1 ECE by 41%-70%, demonstrating substantial gains in probabilistic accuracy and calibration.
CommentsThis manuscript has been submitted to the Knowledge-Based Systems