arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基础适配:面向全场景图像恢复的退化感知可变形分词

Bend the Basics: Degradation-Aware Deformable Tokenization for All-in-One Image Restoration

Zihao He, Yunfeng Wu, Xinchao Wang, Songhua Liu

arXiv 2608.06832首次发表:更新:

AI 中文总结

本文提出FIT模型,通过全流程退化感知设计及任务分词丢弃策略,在五组图像恢复基准上实现最优性能,较现有方法提升0.5~1.1 dB,可适配各类退化图像恢复任务。

AI 中文摘要

全场景图像恢复旨在构建单一模型,以恢复受各类空间非均匀损坏影响的图像。然而,许多统一Transformer依赖固定分块划分:仅在分词后的骨干块中注入任务/退化条件,致使嵌入与重建阶段对局部退化变化不敏感。与现有方法不同,本文提出Flexible Image Transformer(FIT,灵活图像Transformer),其在从分块采样到像素重建的全流程中显式建模退化感知。具体而言,FIT采用轻量退化编码器,从局部退化严重程度预测全局退化向量g与空间退化图M,二者通过自适应变形共同调控分块嵌入与反嵌入。此外,为提升对不同退化类型的鲁棒性,本文引入任务分词丢弃策略,在训练中对任务条件进行正则化。在五个标准基准(BSD68、Rain100L、SOTS、GoPro及LOLv1)上,FIT取得了最优性能:五退化设置下平均峰值信噪比(PSNR)达30.72 dB,三退化设置下达32.83 dB,较近期统一恢复方法提升0.5~1.1 dB;且学习到的偏移量可直接用于可视化退化感知的空间适配。

英文摘要

All-in-one image restoration seeks a single model that can recover images degraded by diverse and spatially non-uniform corruptions. However, many unified Transformers rely on fixed patch partitioning: task/degradation condition is injected only into the backbone blocks after tokenization, leaving the embedding and reconstruction stages insensitive to local degradation variations. In contrast to previous approaches, we present Flexible Image Transformer (FIT) that explicitly models degradation awareness across the entire pipeline, from patch sampling to pixel reconstruction. Specifically, FIT employs a lightweight Degradation Encoder to predict a global degradation vector $\mathbf{g}$ and a spatial degradation map $\mathbf{M}$ from local degradation severity, which jointly condition the patch embedding and unembedding through adaptive deformation. Moreover, to improve robustness across degradation types, we introduce a task-token dropout strategy that regularizes task conditioning during training. On five standard benchmarks (BSD68, Rain100L, SOTS, GoPro, and LOLv1), FIT achieves state-of-the-art performance with 30.72 dB average PSNR on the five-degradation setting and 32.83 dB on the three-degradation setting, outperforming recent unified restoration methods by +0.5$\sim$1.1 dB. Moreover, the learned offsets provide a direct handle for visualizing degradation-aware spatial adaptation.

Comments16 pages, 12 figures. Accepted to ICML 2026

Journal refProceedings of the 43rd International Conference on Machine Learning (ICML 2026), PMLR 306, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑