发表机构
Università di Padova; The University of Queensland(帕多瓦大学; 昆士兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种基于Fisher信息与分类风险几何的标签获取准则,以最小化多类零一分类风险,并开发两阶段自适应算法,实验证明其优于完全分类信息比较器。
AI 中文摘要
我们研究如何分配有限的标注预算以最小化多类别零一分类风险。我们考虑参数化分类问题,其中所有抽样单元的特征可观测,而类别标签可以选择性获取。通过将获取标签提供的Fisher信息与多类别超额风险的局部几何结构相结合,我们推导出一个获取准则,该准则最小化期望多类别超额风险的前导渐近系数。所得规则根据标签信息与扰动活跃贝叶斯决策边界的参数方向的对齐程度来评估标签的价值,而非仅依据后验不确定性或全局参数信息。我们刻画了最优获取设计,确立了其阈值结构,并推导了面向类别面和成本敏感的扩展。一个解析示例表明,后验不确定性与分类价值可能产生不同甚至相反的获取排序。我们进一步开发了一个两阶段自适应程序,在正则条件下达到最优前导风险准则,并为高斯判别分析提供了显式结果。三类别QDA实验展示了所得到的获取几何结构,而对六类别Statlog Landsat卫星数据的应用表明,分类风险获取在实质上不同于基于不确定性的获取和完全分类信息比较器。自适应分类风险设计在所考虑的标注预算范围内比该Fisher比较器获得更低的平均误差,尽管它并未一致优于熵或边际采样,且随着标注预算的增加,各目标策略之间的差异变小。
英文摘要
We study how a limited labeling budget should be allocated to minimize multiclass zero-one classification risk. We consider parametric classification problems in which features are observed for all sampling units while class labels can be acquired selectively. By combining the Fisher information supplied by an acquired label with the local geometry of multiclass excess risk, we derive an acquisition criterion that minimizes the leading asymptotic coefficient of expected multiclass excess risk. The resulting rule values a label according to how strongly its information is aligned with parameter directions that perturb the active Bayes decision boundary, rather than according to posterior uncertainty or global parameter information alone. We characterize the oracle acquisition design, establish its threshold structure, and derive face-specific and cost-sensitive extensions. An analytic example shows that posterior uncertainty and classification value can produce different, and even reversed, acquisition rankings. We further develop a two-stage adaptive procedure that attains the oracle leading-risk criterion under regularity conditions and provide explicit results for Gaussian discriminant analysis. Three-class QDA experiments illustrate the resulting acquisition geometry, while an application to the six-class Statlog Landsat Satellite data shows that classification-risk acquisition can differ materially from both uncertainty-based acquisition and the complete-classification-information comparator. The adaptive classification-risk design attains lower mean error than this Fisher comparator across the labeling budgets considered, although it does not uniformly outperform entropy or margin sampling and differences among the targeted strategies become small as the labeling budget increases.