发表机构
Technological University Dublin(都柏林理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出无标签部署前诊断指标SIM,预测掩蔽令牌剪枝对最差组鲁棒性的影响,并设计批量GPU分割优化,在8个数据集上验证其有效性。
AI 中文摘要
本文证明,基于掩蔽的令牌剪枝是否有助于或损害最差组鲁棒性,可以在部署前无需标签或微调即可预测。对8个虚假相关基准的系统性语义掩蔽研究表明,其对最差组准确率的影响高度不稳定:在某些数据集上相对提升准确率高达82.5%,而在其他数据集上则降低高达100%。我们将这种不稳定性追溯至虚假反转:当虚假属性是背景可分离时,背景补丁获得比真实对象更高的CLIP文本相似度,从而反转了所有文本和注意力引导剪枝方法所依赖的假设。我们引入了虚假反转度量(SIM),一种无标签、部署前的诊断方法,其符号在全部8个数据集上以统计显著性(二项检验p=0.035)预测此效应,并在6种具有清晰前景/背景分割的CLIP架构中保持可靠。朴素掩蔽本身是风险的主要来源:它导致我们评估的所有方法中最大的平均准确率损失,且其自身的逐图像分割步骤是显著的运行时瓶颈。为解决此问题,我们设计了一种批量、无同步的GPU分割例程,将该开销从基线的3.5倍降至1.75倍。根据SIM的符号进行部署门控,可恢复掩蔽的益处同时避免其最严重的失败,在8个数据集中的7个上匹配或超过强剪枝基线。
英文摘要
This paper demonstrate that whether masking-based token pruning helps or hurts worst-group robustness can be predicted before deployment, without labels or fine-tuning. A systematic study of semantic masking across 8 spurious-correlation benchmarks shows its effect on worst-group accuracy is highly unstable: it improves accuracy by up to 82.5\% relative on some datasets and degrades it by up to 100\% on others. We trace this instability to spurious inversion: background patches receive higher CLIP text-similarity than the true object when the spurious attribute is background-separable, inverting the assumption every text- and attention-guided pruning method relies on. We introduce the Spurious Inversion Metric (SIM), a label-free, pre-deployment diagnostic whose sign predicts this effect with statistical significance (binomial $p=0.035$) across all 8 datasets, and remains dependable across 6 CLIP architectures with a clean foreground/background split. Naive masking is itself a major source of risk: it causes the largest average-accuracy loss of any method we evaluate, and its own per-image segmentation step is a significant runtime bottleneck. To address this, we design a batched, synchronization-free GPU segmentation routine that cuts this overhead from 3.5$\times$ to 1.75$\times$ baseline. Gating deployment by SIM's sign recovers masking's benefits while avoiding its worst failures, matching or exceeding a strong pruning baseline on 7 of 8 datasets.
Journal refNeurIPS 2026 Workshop - LIGHT: Deployable Small Foundation Models