arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ProgResViT:用于自适应视觉Transformer的渐进式分辨率与宽度

ProgResViT: Progressive Resolution and Width for Adaptive Vision Transformers

Ali Hojjat, Janek Haberer, Olaf Landsiedel

arXiv 2609.03216首次发表:更新:

发表机构

Kiel University; Hamburg University of Technology (TUHH); UNU-INWEH(基尔大学; 汉堡工业大学; 联合国大学国际水资源研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ProgResViT是一种输入自适应视觉Transformer,通过渐进式多轮推理复用表示并引入PSG机制,在图像分类等任务上实现了更优的精度-计算量权衡。

AI 中文摘要

视觉Transformer(ViTs)通常使用固定的输入分辨率和模型宽度处理每一张图像,而许多图像可以用少得多的计算量完成分类。我们提出ProgResViT,一种输入自适应的ViT,它在多轮中逐步执行推理:第一轮使用低分辨率图像和较窄的子网络处理,当预测足够置信时推理终止;否则模型复用当前轮生成的表示,使用更高分辨率输入和更宽的子网络细化预测。由于所有轮次共享单个骨干网络,我们提出了进度条件软门控(Progress-Conditioned Soft Gating, PSG),它将token融合和层输出条件化于当前轮次、模块块和输入分辨率。在图像分类任务上,将ProgResViT应用于DeiT时,其精度-计算量权衡优于自适应宽度、自适应深度和动态token基线;结合知识蒸馏,基于DeiT的ProgResViT达到84.9%的top-1精度,在可比评估设置下略优于已报道的DeiT-III-S精度。我们还表明,相同设计在自监督DINO表示和下游语义分割任务中也提供了良好的精度-计算量权衡。代码可在该https URL获取。

英文摘要

Vision Transformers (ViTs) typically process every image using a fixed input resolution and model width, even though many images can be classified with substantially less computation. We introduce ProgResViT, an input-adaptive ViT that performs inference progressively across multiple rounds. The first round processes a low-resolution image with a narrow subnetwork. Inference terminates when the prediction is sufficiently confident; otherwise, the model reuses the representations produced in the current round and proceeds with a higher-resolution input and a wider subnetwork to refine its prediction. As all rounds share a single backbone, we propose Progress-Conditioned Soft Gating (PSG), which conditions token fusion and layer outputs on the current round, block, and input resolution. On image classification, applying ProgResViT to DeiT yields better accuracy-compute trade-offs than adaptive-width, adaptive-depth, and dynamic-token baselines. With knowledge distillation, a DeiT-based ProgResViT achieves 84.9% top-1 accuracy, slightly exceeding the reported DeiT-III-S accuracy under a comparable evaluation setting. We show that the same design also provides favorable accuracy-compute trade-offs for self-supervised DINO representations and downstream semantic segmentation. Code is available at https://github.com/ds-kiel/ProgResViT.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑