发表机构
Amazon; The Wharton School, University of Pennsylvania; University of California, Davis(亚马逊; 宾夕法尼亚大学沃顿商学院; 加州大学戴维斯分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对k个序列属性的多校准,建立了匹配的上下样本复杂度界,给出随机学习器,实例化理论为三个典型示例,明确了样本复杂度与群体规模的关系。
AI 中文摘要
校准要求预测器在以自身预测为条件时无偏,多校准则要求在一组群体上同时满足该保证。许多预测任务需要同一条件结果分布的多个相关特征:方差相对于均值定义,偏度相对于均值和方差,条件风险价值相对于分位数。我们研究k个属性序列的多校准,其中每个属性在前面的属性固定后可识别,该框架包含贝叶斯对,但不要求属性源自单一损失。在正则性条件下,对每个固定k≥2,我们建立了对数因子内匹配的上、下样本复杂度界。即使仅存在多项式对数数量的二元群体,达到多校准误差ε也需要\tilde{Ω}(ε^{-(k+2)})个样本;反之,对任意有限群体族\uc0b0\uc0f0,我们给出使用O(ε^{-(k+2)}+ε^{-2}log|\uc0b0\uc0f0})个样本的随机学习器,因此对于多项式规模的群体族,样本复杂度为\tilde{Θ}(ε^{-(k+2)})。我们将该理论实例化为三个典型示例。
英文摘要
Calibration requires a predictor to be unbiased after conditioning on its own predictions. Multicalibration asks for this guarantee simultaneously across a collection of groups. Many prediction tasks ask for several related features of the same conditional outcome distribution: variance is defined relative to the mean, skewness relative to both mean and variance, and conditional value at risk relative to a quantile. We study multicalibration for a sequence of $k$ properties in which each property is identifiable once the preceding properties are fixed. This framework includes Bayes pairs but does not require the properties to arise from a single loss. For every fixed $k\ge2$, we establish matching upper and lower sample-complexity bounds up to logarithmic factors under regularity conditions. Even with only polylogarithmically many binary groups, achieving multicalibration error $\varepsilon$ requires $\widetildeΩ(\varepsilon^{-(k+2)})$ samples. Conversely, for any finite group family $\mathcal G$, we give a randomized learner using $O(\varepsilon^{-(k+2)}+\varepsilon^{-2}\log|\mathcal G|)$ samples. Thus the sample complexity is $\widetildeΘ(\varepsilon^{-(k+2)})$ for polynomial-size group families. We instantiate the theory for three canonical examples.