发表机构
University of Pittsburgh; Xiamen University; Duke University(匹兹堡大学; 厦门大学; 杜克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SimplexUQ 是首个针对单纯形值预测的共形不确定性基准,通过 12 个任务比较现有包装器,揭示全局校准掩盖局部覆盖不均,并强调包装器选择应基于诊断而非通用排名。
AI 中文摘要
共形预测保证了边际覆盖,但单一的校准阈值仍可能使覆盖不均匀分布,过度覆盖简单区域而欠覆盖困难区域。据我们所知,SimplexUQ 是首个用于衡量单纯形值预测中这一分配问题的基准和可复现协议;它比较现有的共形包装器,而非提出新的包装器。其任务套件 SimplexTasks-12 结合了六个受控合成机制与六个冻结预测器的真实任务,涵盖类别概率、主题混合、光谱丰度、细胞类型比例、年龄分布和情感混合。每次比较固定预测器、评分和无响应分层映射,仅改变包装器,并报告边际覆盖、最差分层覆盖、最大差异以及任务内半径和计算量。全局校准可能看似有效但实际表现糟糕:在 CIFAR-10 上,它达到 0.900 的边际覆盖,但在最差熵分层中仅 0.542,而 Mondrian 校准将该分层提升至 0.886,同时将最大差异从 0.358 降至 0.022。然而,没有包装器占主导地位。在平滑的合成异质性下,几种修复方法具有竞争力;固定映射分析表明排名取决于评估组和协议;在 12 任务比较中,Mondrian 在其单一目标分区上对所有 12 个任务具有较低差异,而 BatchMVP 在重叠组上对五个任务具有较低差异。这些是经验比较,而非新的覆盖保证。受控预测器偏差扫描表明,去除预测器偏差仅部分减少全局阈值差异。我们发布任务卡片、结果来源、允许的派生数组和重建说明,并将包装器选择视为诊断性比较而非通用排名。
英文摘要
Conformal prediction guarantees marginal coverage, but a single calibration threshold can still spread that coverage unevenly, over-covering easy regions and under-covering hard ones. SimplexUQ is, to our knowledge, the first benchmark and reproducible protocol for measuring this allocation problem on simplex-valued predictions; it compares existing conformal wrappers rather than proposing a new one. Its task suite, SimplexTasks-12, combines six controlled synthetic regimes with six frozen-predictor real tasks spanning class probabilities, topic mixtures, spectral abundances, cell-type fractions, age distributions, and emotion mixtures. Each comparison fixes the predictor, score, and response-free stratification map, varies only the wrapper, and reports marginal coverage, worst-stratum coverage, max disparity, and within-task radius and compute. Global calibration can look valid while failing badly: on CIFAR-10 it attains 0.900 marginal coverage but only 0.542 in the worst entropy stratum, and Mondrian calibration raises that stratum to 0.886 while reducing max disparity from 0.358 to 0.022. No wrapper dominates, however. Under smooth synthetic heterogeneity, several repairs are competitive; fixed-map analyses show that rankings depend on the evaluation groups and protocol; and in a 12-task comparison, Mondrian has lower disparity on its single target partition for all 12 tasks, whereas BatchMVP has lower disparity over overlapping groups on five. These are empirical comparisons, not new coverage guarantees. A controlled predictor-bias sweep shows that removing predictor bias only partly reduces global-threshold disparity. We release task cards, result provenance, permitted derived arrays, and rebuild instructions, and treat wrapper selection as a diagnostic comparison rather than a universal ranking.
CommentsPreprint. Code and data are available at https://github.com/liangyou03/SimplexUQ and https://huggingface.co/datasets/liangyou03/SimplexTasks-12-data