arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36532stat.MLcs.LG

分层效用校准用于结构化多类决策

Hierarchical Utility Calibration for Structured Multiclass Decisions

  • The University of Osaka(大阪大学)
  • The University of Tokyo(东京大学)
  • RIKEN Center for Advanced Intelligence Project (AIP)(理化学研究所先进智能项目中心(AIP))
  • Mila – Quebec AI Institute(米拉——魁北克人工智能研究所)
  • Université de Montréal(蒙特利尔大学)
  • Nagoya University(名古屋大学)

机构由 AI 辅助整理,请以论文原文为准。

Futoshi Futami, Jerry Huang, Ichiro Takeuchi

AI总结:

针对多类概率预测中的效用校准问题,提出分层效用校准(HUC)方法,通过分解标签树内部节点贡献并逐节点校准,解决了层次结构中效用误差抵消问题,并给出有限样本评估与理论保证。

AI中文摘要:

在多类概率预测中,效用校准(UC)专注于对指定效用的审计,最近作为一种在控制计算和样本需求的同时保证下游决策的方法而受到关注。与此同时,一些多类问题具有有意义的标签层次结构,这些结构在医学和图像分类中发挥重要作用,然而UC如何在层次结构内评估效用仍未被充分理解。我们表明,实现效用与预测平均效用之间的差异可以精确分解为标签树内部节点贡献的总和。这种分解表明,来自不同节点的正负贡献可以相互抵消,并且即使UC很小,层次结构中部分区域剩余的效用误差也可能不小。为解决这一问题,我们提出了分层效用校准(HUC),它在求和之前评估每个节点的贡献,同时保留相同的目标效用、子组和预测效用区间。我们进一步提供了对所有预测效用区间的有限样本评估,并提出了HUC-Boost,该算法仅更新违反约束的内部节点,并为两者提供了理论保证。

英文摘要:

In multiclass probabilistic prediction, Utility Calibration (UC), which focuses auditing on specified utilities, has recently received attention as a way to guarantee downstream decisions while controlling computational and sample requirements. At the same time, some multiclass problems have meaningful label hierarchies that play important roles in medicine and image classification, yet how UC evaluates utility within a hierarchy remains insufficiently understood. We show that the difference between realized utility and predicted mean utility admits an exact decomposition into a sum of contributions from the internal nodes of the label tree. This decomposition shows that positive and negative contributions from different nodes can cancel, and that even when UC is small, the utility errors remaining in parts of the hierarchy need not be small. To address this problem, we propose Hierarchical Utility Calibration (HUC), which evaluates each node contribution before summation while retaining the same target utility, subgroup, and predicted-utility interval. We further provide finite-sample evaluation over all predicted-utility intervals and propose HUC-Boost, which updates only violated internal nodes, with theoretical guarantees for both.

↑