AI 中文总结
该研究针对半监督分类中依赖不确定性的缺失标签,建立似然信息理论,推导费希尔信息分解等,证实此类机制可在固定标记预算下提升分类性能。
AI 中文摘要
在分类任务中,缺失标签通常被视为信息损失的来源。本文研究一种半监督场景,其中标签缺失的概率通过后验分类不确定性依赖于观测特征。在此场景下,缺失指示器不仅是未观测标签的记录,还是与分类器相关机制生成的可观测信号。我们针对这类依赖不确定性的缺失标签建立基于似然的信息理论;在模型设定正确时,推导了费希尔信息分解,将其拆分为部分标记分量与非负机制曲率项;在标签模型和缺失机制联合误设时,得到了对应的Godambe-Eicker-Huber-White敏感性和三明治协方差分解。我们还明确了相关的完整数据基准:有利的缺失性可相对于普通全标记或预算匹配的非信息标记基线增加信息,但无法超过标签和机制指示器均被观测的增强实验中的信息。对于插件分类器,我们将信息分解与基于间隔的超额风险界关联,在常规两分量混合设定中,这产生参数化的\boldsymbol{n^{-1}}超额风险率,常数由判别方向中经干扰调整的信息决定。高斯混合计算和医疗诊断示例表明,在固定标记预算下,依赖不确定性的标记机制可提升估计与分类性能。
英文摘要
Missing labels are usually regarded as a source of information loss in classification. We study a semi-supervised setting in which the probability of label missingness depends on the observed features through posterior classification uncertainty. In this setting, the missingness indicator is not only a record of an unobserved label, but also an observable signal generated by a mechanism linked to the classifier. We develop a likelihood-based information theory for such uncertainty-dependent missing labels. Under correct specification, we derive a Fisher-information decomposition that separates a partial-labeling component from a nonnegative mechanism-curvature term. Under joint misspecification of the label model and the missingness mechanism, we obtain the corresponding Godambe--Eicker--Huber--White sensitivity and sandwich-covariance partitions. We also clarify the relevant complete-data benchmark: favorable missingness can increase information relative to ordinary fully labeled or budget-matched non-informative labeling baselines, but cannot exceed the information in the augmented experiment in which labels and mechanism indicators are both observed. For plug-in classifiers, we connect the information decomposition to margin-based excess-risk bounds. In regular two-component mixture settings this yields the parametric \(n^{-1}\) excess-risk rate, with constants determined by the nuisance-adjusted information in discriminant directions. Gaussian-mixture calculations and a medical diagnosis example illustrate how uncertainty-dependent labeling mechanisms can improve estimation and classification under a fixed labeling budget.
Comments20 pages, 3 figures