共形预测集量化信息增益:一个理论视角
Conformal Prediction Sets Quantify Information Gain: A Theoretical Perspective
浏览论文内容
中文总结 AI 辅助
本文从理论角度证明共形预测集大小减少可作为信息增益度量,引入广义信息度量族,并在11个分类设置中验证其与香农互信息的联系及差异。
中文摘要 AI 辅助
共形预测是一种流行的不确定性量化工具,它输出的预测集具有有限样本覆盖保证。虽然预测集大小通常被用作不确定性的启发式度量,但这种解释的信息论基础仍未被充分理解。在这项工作中,我们通过一种针对集值预测的决策理论化熵推广,为此提供了这样的基础。具体而言,我们引入了一族基于共形预测集大小和覆盖率的广义信息度量。值得注意的是,香农互信息可以用这些度量进行精确的积分表示。随后我们证明,在标准分类设置中,由额外信息引起的共形集大小减少(i)被夹在这一族中依赖于校准的成员之间,并且(ii)满足数据处理不等式,两者均达到有限样本校准和模型误差项。综合来看,我们的结果正式将共形预测与经典信息论量联系起来,并证明将集大小减少作为信息增益度量是合理的。在实证方面,我们在11个分类设置中验证了我们的理论,并表明在贪心特征选择实验中,集大小减少和香农互信息可能对特征进行不同的排序。
英文摘要
Conformal prediction is a popular tool for uncertainty quantification that outputs prediction sets with finite-sample coverage guarantees. While prediction set size is commonly used as a heuristic measure of uncertainty, the information-theoretic basis for this interpretation remains poorly understood. In this work, we provide such a foundation using a decision-theoretic generalization of entropy tailored to set-valued prediction. In particular, we introduce a family of generalized information measures based on the size and coverage of conformal prediction sets. Notably, Shannon mutual information admits an exact integral representation in terms of these measures. We then show that, in standard classification settings, the reduction in conformal set size from additional information (i) is sandwiched between calibration-dependent members of this family and (ii) obeys a data processing inequality, both up to finite-sample calibration and model error terms. Together, our results formally relate conformal prediction to classical information-theoretic quantities and justify using set-size reduction as an information gain metric. Empirically, we validate our theory across 11 classification settings and show that set-size reduction and Shannon mutual information can rank features differently in a greedy feature selection experiment.
发表机构
- MIT(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。