arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Simon-SR:用于文本增强超分辨率的空间自适应调制和视觉提示适配

Simon-SR: Spatially Adaptive Modulation and Visual Prompt Adaptation for Text-Reinforced Super-Resolution

Haotong Cheng, Yuxuan Li, Zijie Cui

arXiv 2607.09351首次发表:更新:

发表机构

College of Electronic Science and Engineering, Jilin University(吉林大学电子科学与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对单图像超分辨率问题,提出Simon-SR框架,利用可学习提示进行语义挖掘和文本-图像融合,结合对比提示学习与空间自适应细化,实验证明该方法超越现有技术,在多项指标上有明显提升。

AI 中文摘要

单图像超分辨率(SISR)从低分辨率输入重建高质量图像。虽近期多模态方法改善了感知质量,但仍对错误先验敏感且需昂贵标注。为解决这些问题,我们提出Simon-SR,一个利用可学习提示进行高效语义挖掘和稳健文本-图像融合的多模态SISR框架。我们的方法将对比提示学习与提示引导的空间自适应细化相结合以增强多模态对齐。实验表明Simon-SR超越了现有方法,在PSNR、SSIM和LPIPS上有显著提升。代码将发布。

英文摘要

Severe downsampling makes single image super-resolution (SISR) an ill-posed problem, in which the language-guided multi-modal methods are especially vulnerable to erroneous priors and require costly textual annotations. To address these issues, we propose Simon-SR, a multi-modal SISR framework leveraging learnable prompts for efficient semantic mining and robust text-image fusion. Simon-SR treats textual semantics as learnable latent variables. Specifically, the Contrastive Prompt Learning (CPL) mines instance-level semantics from unannotated images with frozen CLIP encoders, and Prompt-Guided Spatially Adaptive Refinement (PSAR) injects them through attention-gated multi-modal fusion. On CUB and COCO2017 at $\times4$ and $\times16$, Simon-SR gains up to 0.50 dB PSNR and 0.0133 SSIM over current SOTA. The project is available at https://github.com/CHT05017/ICAIS26-SimonSR

CommentsMulti-modal Single Image Super-Resolution

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑