基于威布尔混合模型中信息性缺失标签的半监督分类
Semi-Supervised Classification with Informative Missing Labels in Weibull Mixture Models
- The University of Queensland(昆士兰大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对威布尔混合模型下的部分标注样本,构建依赖特征的随机缺失标签机制,推导分类器的费希尔信息与误差率展开式,通过数值及半合成分析验证建模缺失标签可降低分类误差、改进决策边界估计。
AI中文摘要:
我们研究了由两分量威布尔混合模型产生的部分标注样本的半监督分类问题。所有数据均可观测到特征,但部分类别标签存在缺失。标签缺失概率被建模为分类不确定性的函数,形成了与威布尔混合分类器共享参数的、依赖于特征的随机缺失(MAR)机制。因此,除观测特征和可用类别标签外,缺失标签指示符还能提供关于分类器的额外信息。在公共威布尔形状参数下,贝叶斯规则最多有一个正决策边界,且当该规则为非恒定函数时,此边界唯一;在形状参数不等的情况下,贝叶斯规则可拥有两个正决策边界。我们对这些决策区域进行了刻画,推导了缺失性模型中冗余参数调整后分类器的费希尔信息,并得到了插件样本规则相对于贝叶斯误差的期望误差率的决策边界展开式。该展开式为单边界和双边界情况提供了分类特定的渐近相对效率公式,表明费希尔信息的正定增长是一阶期望误差率减小的充分非必要条件。数值研究及基于硬盘故障数据的半合成分析表明,对依赖于特征的标签缺失性进行建模可降低期望误差率,并改进决策边界估计。
英文摘要:
We consider semi-supervised classification from a partially classified sample arising from a two-component Weibull mixture. The feature is observed for all data, whereas some class labels are missing. The probability of a missing label is modelled as a function of classification uncertainty, giving a feature-dependent missing-at-random (MAR) mechanism that shares parameters with the Weibull-mixture classifier. The missing-label indicators can therefore provide information about the classifier in addition to the observed features and available class labels. Under a common Weibull shape, a Bayes' rule has at most one positive decision boundary, which is unique when the rule is nonconstant; under unequal shapes, it can have two. We characterise these decision regions, derive the Fisher information for the classifier after adjustment for nuisance parameters in the missingness model, and obtain a decision-boundary expansion of the expected error rate of the plug-in sample rule relative to the Bayes error. The expansion yields classification-specific asymptotic relative efficiency formulas for the one- and two-boundary cases and shows that a positive-definite increase in Fisher information is sufficient, but not necessary, for a smaller first-order expected error rate. Numerical studies and a semi-synthetic analysis based on hard-drive failure data illustrate potential reductions in expected error rate and improvements in decision-boundary estimation from modelling feature-dependent label missingness.