arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22843stat.COmath.STstat.TH

指数混合模型半监督分类中的有利缺失机制

Favourable Missingness in Semi-Supervised Classification for Exponential Mixture Models

Huanchao Zhou, Jinran Wu, Fariborz Setoudehtazang, Geoffrey J. McLachlan

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对指数混合模型,研究标签缺失概率依赖分类不确定性的半监督分类机制,推导贝叶斯规则与费舍尔信息分解,通过数值与蒙特卡洛实验验证该机制的渐近及有限样本性能。

中文摘要 AI 辅助

半监督分类器通常从所有特征可观测但部分类别标签缺失的样本中训练。当标签缺失与观测数据无关时,不可用的类别成员会相对于完全标注样本降低费舍尔信息。本文研究不同机制:标签缺失概率依赖于后验分类不确定性,因此观测到的缺失标签指示符本身可携带关于贝叶斯决策边界的信息。基于Ahfock和McLachlan的条件加权信息分解,本文针对两分量指数混合模型发展该现象。尽管指数模型非高斯、不对称且支撑于正半轴,其后验对数几率仍线性于特征。本文推导贝叶斯规则及其精确错误率,构建熵-逻辑斯蒂与平方判别缺失机制,得到完整部分标注似然;将费舍尔信息分解为完整数据信息、因缺失标签产生的条件加权损失、缺失标签贡献的信息。数值积分确定完整似然分类器渐近相对效率高于或低于1的区域;有限训练样本的蒙特卡洛实验大致支持总体计算,渐近预测的最大偏差发生在相对效率跨越1的过渡附近。

英文摘要

Semi-supervised classifiers are commonly trained from samples in which all features are observed but some class labels are missing. When label missingness is independent of the observed data, unavailable class memberships reduce Fisher information relative to a completely classified sample. We study a different regime in which the probability of label missingness depends on posterior classification uncertainty, so that the observed missing-label indicators can themselves carry information about the Bayes decision boundary. Building on the conditionally weighted information decomposition of Ahfock and McLachlan, we develop this phenomenon for a two-component exponential mixture. Although the exponential model is non-Gaussian, asymmetric, and supported on the positive half-line, its log-posterior odds remain linear in the feature. We derive Bayes' rule and its exact error rate, formulate entropy-logistic and squared-discriminant missingness mechanisms, and obtain the full partially classified likelihood. We then derive a decomposition of the Fisher information into the complete-data information, the conditionally weighted loss due to missing labels, and the information contributed by the missing labels. Numerical quadrature identifies regions in which the full likelihood classifier has asymptotic relative efficiency above or below one. Monte Carlo experiments with finite training samples broadly support the population calculations, with the largest departures from the asymptotic predictions occurring near the transition at which the relative efficiency crosses one.

补充信息

↑