发表机构
RWTH Aachen University; Institute of Statistics and AI Center(亚琛工业大学; 统计与人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对分类器条件分布缺乏代表性的问题,提出上下文自适应阈值化方法,通过调整阈值保持误报率并提升代表性,经非参数估计实现,在FICOS信用评分上媲美最先进方法。
AI 中文摘要
通常,分类器和监测程序是通过优化诸如误分类率等目标函数,从标记数据中训练得到的。这可能导致在给定重要外部变量的条件下,结果(即标签)的条件分布缺乏代表性,与总体中的条件分布律不同。我们展示了如何修改任意给定的阈值型分类器或监测规则,通过将阈值适应于协变量$Z$(即上下文)来分配灵敏度,同时保持误报率,从而实现具有代表性的条件标签预测。当报警事件未知时,该方法还允许(近似地)以阈值规则的形式推断该事件。该方法通过一种计算成本低廉的非参数估计程序实现,并利用非渐近误差界和渐近分布理论(包括经验过程理论)研究了其性质。这些结果使得构建一致置信带、函数性假设检验和变化检测程序成为可能。对于可解释机器学习中常用的著名FICOS信用评分示例,阈值自适应产生了一个易于解释的决策规则,在常见分类指标上可与包括Transformer在内的最先进方法相竞争。
英文摘要
Commonly, classifiers and monitoring procedures are trained from labeled data by optimizing an objective such as the misclassification rate. This may lead to unrepresentative conditional distributions of the outcome (the labels) given important external variables, different from the conditional laws in the population. We show how to modify any given threshold-type classifier resp. monitoring rule to achieve representative conditional label prediction by using adapting the threshold to a covariate $Z$ (the context) to distribute sensitivity while maintaining the false alarm rate. In case that the alarm event is unknown, this approach also allows to (approximately) infer the event in terms of a thresholding rule. The approach is implemented by a computationally cheap nonparametric estimation procedure, and its properties are studied in terms of nonasymptotic error bounds and asymptotic distribution theory including empirical process theory. These results allow to construct uniform confidence bands, functional hypothesis tests and change-detection procedures. For the well known FICOS credit scoring example, often used in interpretable machine learning, threshold adaptation leads to an easily interpretable decision rule which can compete with state of the art methods including transformers, in terms of common classification metrics.