发表机构
Independent Researcher(独立研究者)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对基于ViT的图像超分辨率模型计算复杂度高的问题,提出聚类单元级相似性变换器(CUST)。该模型整合全局与局部信息,用重叠窗口捕局部依赖、算残差提高频细节,在计算效率和恢复性能间达平衡,内存占用低、推理速度快。
AI 中文摘要
最近,基于视觉变换器(ViT)的模型在图像超分辨率方面表现出卓越性能。然而,ViT在空间分辨率上的二次计算复杂度严重限制了其效率,导致高延迟和大量内存消耗。为缓解此问题,已提出各种基于窗口的注意力机制,但它们在一定程度上损害了ViT的主要优势——长距离依赖建模。为克服这些限制,我们提出了聚类单元级相似性变换器(CUST),这是一种有效整合全局和局部信息的新型架构。具体而言,CUST使每个补丁能够在其局部窗口之外的扩展区域范围内聚合并关注相似补丁,从而获得广泛的上下文理解。此外,它采用重叠注意力窗口来捕捉局部依赖关系,同时通过计算原始特征与其下采样-上采样对应特征之间的残差差异来明确提取高频细节。综合实验表明,我们提出的模型在计算效率和恢复性能之间实现了实际平衡。在实际约束下,与最近的全局上下文或轻量级模型相比,它具有更低的内存占用和更快的推理速度。代码可在[此https URL]获取。
英文摘要
Recently, Vision Transformer (ViT)-based models have exhibited remarkable performance in image super-resolution. However, the quadratic computational complexity of ViTs with respect to spatial resolution severely constrains their efficiency, leading to high latency and massive memory consumption. To alleviate this, various window-based attention mechanisms have been proposed; yet, they inherently compromise the long-range dependency modeling that is the primary advantage of ViTs. To overcome these limitations, we propose the Clustered Unit-level Similarity Transformer (CUST), a novel architecture that efficiently integrates global and local information. Specifically, CUST enables each patch to aggregate and attend to similar patches within a broadened regional scope outside its local window, thereby capturing extensive contextual understanding. Furthermore, it employs overlapping attention windows to capture local dependencies, while explicitly extracting high-frequency details by computing the residual difference between the original features and their downsampled-upsampled counterparts. Comprehensive experiments demonstrate that our proposed model achieves a practical balance between computational efficiency and restoration performance. It achieves a lower memory footprint and faster inference speed compared to recent global context or lightweight models under realistic constraints. Code is available at [https://github.com/jwgdmkj/CUST].
Comments15 pages, 7 figures