CORE-STACK+:深度堆叠泛化的元学习
CORE-STACK+: Meta-Learning for Deep Stacked Generalization
- Istanbul Technical University(伊斯坦布尔理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对异构视觉骨干堆叠中的多重共线性与校准崩溃问题,提出CORE-STACK+预处理流程,含核冗余过滤、元特征门、谱自适应岭惩罚和贝叶斯混合器,在六个基准上显著提升精度与校准并降低计算开销。
AI中文摘要:
堆叠异构视觉骨干网络(CNN、ViT及混合模型)是提升准确率、校准性和鲁棒性的标准做法,但两种相互关联的病理限制了其收益。预测空间中的多重共线性使元学习器的Gram矩阵病态化,导致权重方差膨胀,并在狭窄流形上产生脆弱解。校准崩溃通过朴素线性堆叠加剧了各组成模型的校准误差,因此增加更多模型可能反而损害期望校准误差(ECE)。现有解决方案(如岭正则化、贪心选择、模型汤和SWAG)最多只能解决其中一个问题,且均未能在异构预测池中联合处理条件化和校准问题。我们提出CORE-STACK+,一种包含四个组件的预处理流程:(i)基于核的冗余过滤器,利用中心核对齐(CKA)[23]移除皮尔逊相关性无法发现的非线性模型间依赖;(ii)一个参数少于15K的可微分元特征门,学习对集成统计量的逐样本注意力;(iii)基于Marchenko-Pastur信噪分解推导的自适应谱岭惩罚λ*=lmax(Chat)/SNR(Chat),消除了嵌套交叉验证;(iv)一个拉普拉斯近似贝叶斯混合器,替代逆RMSE启发式方法。我们证明了一个PAC-Bayes超额风险界,首次联合考虑了预测空间冗余和元学习器容量。在六个基准上,CORE-STACK+在ImageNet-1K上top-1准确率提升+1.8%,在ImageNet-C上mCE降低-4.2,在ADE20K上mIoU提升+0.9,在COCO上AP提升+1.3,同时将保留模型数量减少35-57%,推理FLOPs降低最多41%。ECE比深度集成改善2.1倍,且无需事后温度缩放。
英文摘要:
Stacking heterogeneous vision backbones (CNNs, ViTs, and hybrids) is the de facto recipe for accuracy, calibration, and robustness, yet two coupled pathologies limit its returns. Prediction-space multicollinearity ill-conditions the meta-learner's Gram matrix, inflating weight variance and producing brittle solutions on a thin manifold. Calibration collapse compounds constituent miscalibration through naive linear stacking, so adding more models can hurt expected calibration error (ECE). Existing remedies, ridge regularization, greedy selection, model soups, and SWAG address at most one of these issues, and none jointly target conditioning and calibration in heterogeneous prediction pools. We introduce CORE-STACK+, a preconditioning pipeline with four components: (i) a kernelized redundancy filter that removes non-linear inter-model dependencies invisible to Pearson correlation, using Centered Kernel Alignment (CKA) [23]; (ii) a $<15$K-parameter differentiable meta-feature gate that learns per-sample attention over ensemble statistics; (iii) a spectrum-adaptive Ridge penalty $lambda^{star}=lmax(Chat)/SNR(Chat)$ derived from a Marchenko-Pastur signal-noise decomposition, eliminating nested cross-validation; and (iv) a Laplace-approximate Bayesian blender replacing inverse-RMSE heuristics. We prove a PAC-Bayes excess-risk bound that, for the first time, jointly accounts for prediction-space redundancy and meta-learner capacity. Across six benchmarks, CORE-STACK+ delivers $+1.8\%$ top-1 on ImageNet-1K, $-4.2$ mCE on ImageNet-C, $+0.9$ mIoU on ADE20K, and $+1.3$ AP on COCO, while reducing retained models by 35-57% and inference FLOPs by up to $41%$. ECE improves $2.1\times$ over deep ensembles without post hoc temperature scaling.