arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17966cs.CV

SFMformer:用于轻量级图像超分辨率的空间-频率调制Transformer

SFMformer: A Spatial-Frequency Modulation Transformer for Lightweight Image Super-Resolution

发表机构淡江大学
查看机构详情
  • Tamkang University(淡江大学)

机构由 AI 辅助整理,请以论文原文为准。

Chih-Hsiang Yang, Chia-Min Lin, Ching-Yu Tsai, Yung-Che Wang, Jen-Shiun Chiang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对轻量级图像超分辨率提出SFMformer,通过分离注意力的选择与聚合质量设计双模块,在低参数量下实现多数基准指标最优,且在树莓派5上验证了实用性。

中文摘要 AI 辅助

稀疏注意力机制会对所有token对打分,但仅传播最强的连接,目前是轻量级图像超分辨率领域最高效Transformer的核心支撑。本文发现,稀疏化改变了改进这类网络的核心要义:密集注意力层仅在被关注特征的聚合处对表示质量有要求;而稀疏层有两处要求,因为top-k算子先决定哪些token保留,再处理这些token,在选择阶段被丢弃的token无法在下游恢复,因此选择质量与聚合质量是可分离的目标,分别由注意力前后的模块处理。我们通过将渐进式聚焦注意力的输入搭配双分支空间增强,输出搭配小波域调制,形成SFMformer来验证这一设计。在全部15组基准尺度对上分别及联合测试各模块,发现它们的增益并非可加:9组联合增益超过各模块单独增益之和,差异符号由较弱模块自身贡献量预测(相关系数r=-0.72),因此二者在缓解不同约束时复合生效,缓解相同约束时重叠生效。每个块而非每层启用一次频谱调制,以约六分之一的成本保留效果,使模型在所有尺度上参数均低于100万。SFMformer在5个基准、3个上采样因子的30项PSNR/SSIM指标中28项排名第一。我们报告了该配对设计无效的案例,并在树莓派5(Raspberry Pi 5)上部署模型,确认其在严格资源预算下具备实用性。

英文摘要

Sparse attention mechanisms, which score all token pairs but propagate only the strongest, now underpin the most efficient Transformers for lightweight image super-resolution. This paper observes that sparsification changes what it means to improve such a network. A dense attention layer has one place where representation quality matters: the aggregation of attended features. A sparse layer has two, because the top-k operator first decides which tokens survive and only then decides what to do with them, and a token discarded at the selection stage cannot be recovered downstream. Selection quality and aggregation quality are therefore separable targets, addressed by modules placed before and after the attention respectively. We test this by pairing a dual-branch spatial enhancement on the input of a progressive focused attention with a wavelet-domain modulation on its output, forming SFMformer. Measuring each module alone and jointly over all fifteen benchmark-scale pairs, we find their gains are not additive: the joint gain exceeds the sum of the individual gains on nine pairs, and the sign of the discrepancy is predicted by how much the weaker module contributes on its own (r = -0.72), so the two compound when they relieve different constraints and overlap when they relieve the same one. Enabling spectral modulation once per block rather than once per layer retains the effect at roughly one-sixth of its cost, keeping the model below one million parameters at every scale. SFMformer ranks first on 28 of 30 PSNR/SSIM entries across five benchmarks and three upscaling factors. We report the cases where the pairing does not help, and deploy the model on a Raspberry Pi 5 to confirm the design is practical under tight resource budgets.

补充信息

↑