扩展全共形图像分类器
Scaling Full Conformal Image Classifiers
浏览论文内容
中文总结 AI 辅助
针对全共形预测计算成本高的问题,提出目标全共形预测(T-FCP),利用零样本视觉语言模型剪枝标签,仅对候选子集应用FCP,在保持覆盖保证的同时降低开销,并在ImageNet等基准上实现高效稳定的预测集。
中文摘要 AI 辅助
共形预测提供了具有分布无关覆盖保证的集合值预测,使其在高风险图像分类中具有吸引力。然而,分割共形预测数据效率低下,而全共形预测(FCP)尽管具有更强的统计效率,但在大规模应用中计算成本过高,因为它在测试时需要针对每个候选重新拟合模型。我们通过利用零样本视觉语言模型(VLM)来指导大规模标签空间中的可扩展FCP,解决了这一限制。我们引入了目标全共形预测(T-FCP),它使用轻量级归纳共形预测器来剪除不太可能的标签,并仅对剩余候选应用FCP,从而在保留组合共形过程正式保证的同时减少计算量。我们进一步提出了稳定化在线LDA(SO-LDA),一种基于秩一逆协方差更新的高效VLM适应求解器。在包括ImageNet在内的多个基准测试中,T-FCP实现了实用的全共形图像分类,测试时开销适中,产生高效的预测集,并且比分割共形替代方案具有更稳定的经验覆盖率。
英文摘要
Conformal prediction provides set-valued predictions with distribution-free coverage guarantees, making it attractive for high-stakes image classification. However, split conformal prediction is data-inefficient, while full conformal prediction (FCP), despite its stronger statistical efficiency, is computationally prohibitive at scale because it requires candidate-specific model refits at test time. We address this limitation by leveraging zero-shot vision-language models (VLMs) to guide scalable FCP in large label spaces. We introduce Targeted Full Conformal Prediction (T-FCP), which uses a lightweight inductive conformal predictor to prune unlikely labels and applies FCP only to the remaining candidates, reducing computation while retaining the formal guarantee of the combined conformal procedure. We further propose Stabilized Online LDA (SO-LDA), an efficient VLM adaptation solver based on rank-one inverse-covariance updates. Across multiple benchmarks, including ImageNet, T-FCP enables practical full-conformal image classification with modest test-time overhead, yielding efficient prediction sets and more stable empirical coverage than split conformal alternatives.
发表机构
- ETH Zürich(苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。