发表机构
University of Technology Graz; Know-Center Research GmbH; Institute of Visual Computing(格拉茨技术大学; 诺中心研究有限公司; 视觉计算研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出改进与剪枝(I&P)方法,将迭代幅度剪枝融入主动学习重训练循环,以几乎零额外成本获得稀疏中奖彩票,在高达95%稀疏度下保持准确率,并缓解深度主动学习的计算瓶颈。
AI 中文摘要
彩票假设(lottery ticket hypothesis)认为存在“中奖彩票”:即稀疏子网络,当从原始初始化状态单独训练时,其准确率能与完整稠密网络相匹配。发现此类彩票的主要方法是迭代幅度剪枝(iterative magnitude pruning),该方法在每一轮中交替进行剪枝和从头开始的完整重训练,直至收敛,并重复多个周期。类似地,深度主动学习(deep active learning)在每一轮获取新标签后,也会从头开始重新训练模型。尽管这两种范式都依赖于迭代重训练并带来大量计算开销,但它们一直被分开研究。我们观察到,基于池的主动学习(pool-based active learning)固有的迭代训练循环恰好提供了迭代幅度剪枝所利用的精确计算结构,并据此提出了“改进与剪枝”(Improve & Prune, I&P)方法,该方法将幅度剪枝集成到每个主动学习重训练周期中,且几乎不增加额外成本。这引出了一个关键的实证问题:在主动学习的非平稳数据机制下,迭代幅度剪枝能否产生中奖彩票?我们跨多种采集函数(acquisition functions)、架构家族和图像分类数据集(包括主动微调场景)对此问题进行了研究。我们的结果表明,I&P 在每个主动学习迭代中都能产生稀疏、可部署的模型。这些模型在稀疏度高达95%时,其准确率与稠密对应模型相匹配,从而有效地将中奖彩票作为主动学习流程的副产品获得。这些逐迭代的稀疏模型可以解决两个计算瓶颈——每轮模型重训练和未标注池上的采集评分——目前这两个瓶颈阻碍了深度主动学习(DAL)在大规模架构和大规模未标注池上的实际应用。
英文摘要
The lottery ticket hypothesis posits the existence of winning tickets: sparse subnetworks that, when trained in isolation from their original initialization, match the accuracy of the full dense network. The predominant method for discovering such tickets, iterative magnitude pruning, alternates pruning with full retraining from scratch until convergence over many cycles. Similarly, deep active learning also retrains a model from scratch after each acquisition round as new labels become available. Despite this shared reliance on iterative retraining with a substantial computational overhead, the two paradigms have been studied separately. We observe that the iterative training loop inherent to pool-based active learning already provides the exact computational structure that iterative magnitude pruning exploits, and propose Improve & Prune (I&P), a method that integrates magnitude pruning into each active learning retraining cycle at practically no additional cost. This raises a key empirical question: can iterative magnitude pruning produce winning tickets under the non-stationary data regime of active learning? We investigate this question across multiple acquisition functions, architecture families, and image classification datasets, including an active fine-tuning scenario. Our results demonstrate that I&P yields sparse, deployable models at each active learning iteration. Those match the accuracy of their dense counterparts at sparsities up to 95%, effectively obtaining winning tickets as a byproduct of the active learning pipeline. These per-iteration sparse models can address two computational bottlenecks - per-round model retraining and acquisition scoring over the unlabeled pool - that currently prevent the practical adoption of DAL on large architectures and large unlabeled pools.