发表机构
Institute for Cancer Genetics and Informatics Oslo University Hospital; Department of Informatics University of Oslo; SFI Visual Intelligence UiT The Arctic University of Norway(奥斯陆大学医院癌症遗传学与信息学研究所; 奥斯陆大学信息学系; 挪威北极大学视觉智能SFI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对视觉 Transformer 的深度计算冗余,提出 TWT 事后方法,融合冗余层为单一替代层,减参降算,半深度下保持竞争力并提升病理学任务表现。
AI 中文摘要
近期研究结果表明,视觉 Transformer 会收敛到局部相似的计算阶段,这意味着存在一定程度的深度计算冗余。然而,现有利用这种冗余的方法要么无法减少推理计算量,要么严重降低模型表达能力。在本工作中,我们形式化了一个统一的块冗余视图,将几何结构与特定的替代干预措施解耦。随后,我们提出了 Transformer-Within-Transformer (TWT),这是一种事后方法,将连续的冗余层组融合为单个学习到的替代层。TWT 减少了参数数量和推理计算量,同时在自然图像上使用一半深度仍能与原始模型保持竞争力,并且在多个下游组织病理学设置中,TWT 匹配甚至改进了原始基线。
英文摘要
Recent findings indicate that Vision Transformers settle into locally similar computational phases, implying a level of depthwise computational redundancy. However, existing methods to exploit this redundancy either fail to reduce inference compute or severely degrade model expressivity. In this work, we formalise a unified view of block redundancy that decouples the geometry from specific surrogate interventions. We then introduce Transformer-Within-Transformer (TWT), a post-hoc method that fuses contiguous groups of redundant layers into a single learned surrogate layer. TWT reduces parameter count and inference compute while remaining competitive with original models using half the depth on natural images, and in several downstream histopathology settings, TWT matches or even improves on the original baseline.
Comments22 pages, 6 figures, 6 tables. Accepted at the 37th British Machine Vision Conference (BMVC 2026)