发表机构
Iowa State University; Independent University of Bangladesh(爱荷华州立大学; 孟加拉独立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出保留校准的剪枝(CPP)方法,通过拆分共形预测保证剪枝后模型的覆盖率,在DBpedia-14等数据集上实现更小预测集尺寸,提升多数任务的预测效率。
AI 中文摘要
一旦剪枝模型独立于共形校准拆分被固定,拆分共形预测(Split conformal prediction,而非剪枝规则)可提供有限样本边际覆盖率。本文研究单独效率问题:剪枝能否足够好地保留得分几何,以获得更小的有效预测集?保留校准的剪枝(Calibration-Preserving Pruning, CPP)在基础剪枝得分中加入非一致性梯度显著性,并使用互不相交的剪枝、验证选择、共形校准及测试拆分。有界得分扰动意味着有界共形分位数偏移和可控集合膨胀,但未使通用覆盖率定理成为CPP特有的。50%稀疏度下五次随机种子的Qwen2.5-1.5B结果显示,在大标签任务上增益最大。在DBpedia-14上,CPP-SparseGPT将平均集合大小从10.1降至8.6,同时准确率从0.347变为0.366;CPP-Wanda将平均集合大小从11.2降至9.0,准确率从0.310变为0.295。在15个数据集-稀疏度组合中,CPP-SparseGPT在13个组合中产生更小集合,在11个组合中准确率更高。匹配对照实验表明,通用有监督梯度可解释大部分增益:真实标签CPP与匹配的Wanda+SNIP无统计学差异,而阈值感知候选标签CPP在显式准确率和离线计算成本下达到7.8的平均集合大小。RoBERTa-base和Llama-3-8B的诊断结果支持可迁移性,但本文结论仍限于对可靠性敏感的分类任务。
英文摘要
Split conformal prediction, not the pruning rule, supplies finite-sample marginal coverage once a pruned model is fixed independently of the conformal calibration split. We study the separate efficiency problem: can pruning preserve score geometry well enough to obtain smaller valid prediction sets? Calibration-Preserving Pruning (CPP) augments a base pruning score with nonconformity-gradient saliency and uses disjoint pruning, validation-selection, conformal-calibration, and test splits. Bounded score perturbations imply bounded conformal-quantile shifts and controlled set inflation, but do not make the generic coverage theorem CPP-specific. Final five-seed Qwen2.5-1.5B results at 50\% sparsity show the largest gains on large-label tasks. On DBpedia-14, CPP-SparseGPT reduces mean set size from \(10.1\) to \(8.6\) while changing accuracy from \(0.347\) to \(0.366\); CPP-Wanda reduces \(11.2\) to \(9.0\) with an accuracy trade-off from \(0.310\) to \(0.295\). Across 15 dataset--sparsity cells, CPP-SparseGPT produces smaller sets in 13 and higher accuracy in 11. Matched controls show that generic supervised gradients explain much of the gain: true-label CPP is not statistically resolved from matched Wanda+SNIP, whereas threshold-aware candidate-label CPP reaches \(7.8\) mean set size at explicit accuracy and offline-compute costs. RoBERTa-base and Llama-3-8B diagnostics support transfer, but our claims remain limited to reliability-sensitive classification.