arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10989cs.CVcs.AI

发挥寄存器的作用:用于视觉Transformer中令牌剪枝的任务寄存器

Putting Registers to Work: Task Registers for Token Pruning in Vision Transformers

Hongsen Cao, Mona Jaber, Shanxin Yuan, Ahmed Sayed

AI总结:

本文探究视觉Transformer令牌剪枝策略在不同任务间的迁移性,提出任务自适应剪枝TAP,引入任务寄存器,在ρ=0.5时实现ADE20K、COCO上的性能提升且保持ImageNet-1K竞争力。

AI中文摘要:

令牌剪枝策略通常是为单一识别流水线设计的,但预训练的视觉Transformer(Vision Transformers)会被用于空间需求不同的各类任务。本文探究了剪枝策略的哪些部分可在图像分类、语义分割和目标检测间迁移。针对每个流水线,受控探测会冻结无剪枝的检查点,在单个符合条件的层上应用一系列无参数的缩减准则,且无需重新训练。探测结果揭示了三处差异:分割和检测对准则的排序不同;分类对最早层中基于注意力的剪枝尤为敏感;密集任务偏好相反的恢复端点。这些发现启发了任务自适应剪枝(Task-Adaptive Pruning, TAP)的提出。现有的寄存器令牌可作为任务无关的特征伪影存储,而TAP则为每个任务引入一个任务寄存器,仅激活当前任务的寄存器。其演化状态会对令牌排序、在网络深度上分配精确的移除预算,并为密集特征设置恢复尺度。在最终保留率ρ=0.5时,本文联合适配的模型TAP-J在ADE20K数据集上达到47.0的平均交并比(mIoU),编码器吞吐量为1.30倍;在COCO数据集上达到53.7的框平均精度(box AP),编码器吞吐量为1.32倍,同时在ImageNet-1K上仍保持竞争力。

英文摘要:

Token-pruning policies are usually designed for a single recognition pipeline, but pretrained Vision Transformers are reused across tasks with different spatial demands. We ask which parts of a pruning policy transfer across image classification, semantic segmentation, and object detection. For each pipeline, controlled probes freeze the no-pruning checkpoint and apply a series of parameter-free reduction criteria at one eligible layer at a time without retraining. The probes reveal three differences: segmentation and detection rank the criteria differently, classification is especially sensitive to attention-based pruning in the earliest layers, and the dense tasks prefer opposite recovery endpoints. These findings motivate Task-Adaptive Pruning (TAP). Existing register tokens serve as task-agnostic storage for feature artifacts. TAP instead introduces one task register per task and activates only the current one. Its evolving state ranks tokens, distributes an exact removal budget over depth, and sets the recovery scale for dense features. At a final keep rate of $ρ=0.5$, our jointly adapted model, TAP-J, reaches $47.0$ mIoU at $1.30\times$ encoder throughput on ADE20K and $53.7$ box AP at $1.32\times$ encoder throughput on COCO while remaining competitive on ImageNet-1K.

补充信息

↑