arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

循环视觉Transformer的训练十字路口:循环、神经ODE与深度监督

Training Crossroads for Recurrent Vision Transformers: Recurrence, Neural ODEs, and Deep Supervision

Grzegorz Gruszczynski, Pawel Olszowiec, Michal Byra, Grzegorz Stefanski, Alberto Presta

arXiv 2608.04879首次发表:更新:

AI 中文总结

本文针对单模块循环视觉Transformer,在CIFAR-100上对比三种训练推理模式,明确了循环与标准ViT的适用约束、ODE求解器阶数的作用,以及深度监督对训练时域外鲁棒性的提升效果。

AI 中文摘要

视觉Transformer(ViTs)在图像识别任务中表现出色,但若每个模块都独立参数化,其参数量会随深度线性增长。单模块循环ViT(bViT)通过重复使用同一个共享模块消除了这种参数量增长。本文并未提出新架构,而是固定一个bViT,在统一的CIFAR-100实验协议下,对三种训练与推理模式开展受控实证研究,旨在回答三个问题:(i)在FLOPs匹配或参数内存匹配的条件下,循环机制何时优于独立参数化的深度结构?(ii)当残差循环模块通过ODE求解器训练时,求解器阶数是起到数值细化作用,还是作为一种架构偏置?(iii)训练时域之外的鲁棒性需要在标称精度上付出多少代价?研究发现,当FLOPs是主要约束时,标准ViT仍是更优选择;而在内存约束下,循环ViT能提供更好的精度-参数权衡。与残差网络作为ODE欧拉离散化的经典观点一致,残差循环模块的连续时间对应形式是状态减去向量场$\boldsymbol{\dot{z}}=F_\theta(z)-z$;尽管这一区别原理上已知,但当模块被封装为黑盒向量场时很容易被违背,本文量化了其在少量精度点上的代价。由于向量场与求解器是联合学习的,高阶求解器起到的是求解器诱导的架构偏置作用,而非提升数值精度,且其带来的增益并不均匀。最后,分阶段深度监督勾勒出一条精度-鲁棒性边界:它不会提升标称精度,但在远超训练时域的情况下性能仍能平缓下降,而朴素循环结构此时会崩溃至接近随机的性能。

英文摘要

Vision Transformers (ViTs) achieve strong image-recognition performance, but their parameter count grows linearly with depth when each block is independently parameterized. Single-block recurrent ViTs (bViT) remove this growth by repeatedly applying one shared block. Rather than proposing a new architecture, we fix a bViT and provide a controlled empirical characterization of three training and inference regimes under a common CIFAR-100 protocol, asking: (i)~when does recurrence beat independently parameterized depth---at matched FLOPs or at matched parameter memory? (ii)~when a residual recurrent block is trained through an ODE solver, does solver order act as numerical refinement or as an architectural bias? and (iii)~what does robustness beyond the training horizon cost in nominal accuracy? We find that standard ViTs remain preferable when FLOPs are the primary constraint, whereas recurrent ViTs offer a better accuracy--parameter trade-off under memory constraints. Consistent with the standard view of residual networks as Euler discretizations of ODEs, the continuous-time analogue of a residual recurrent block is the state-subtracted vector field $\dot{z}=F_θ(z)-z$; although known in principle, this distinction is easy to violate when the block is wrapped as a black-box vector field, and we qualify the cost at few accuracy points. Because the vector field is learned jointly with the solver, higher-order solvers act as a solver-induced architectural bias rather than a numerical-accuracy improvement, and their gains are not uniform. Finally, stage-wise deep supervision traces an accuracy--robustness frontier: it does not improve nominal accuracy, but degrades gracefully far beyond the training horizon, where naive recurrence collapses to near-random performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑