TOLA:面向扩散式文本图像超分辨率的文本感知单步潜在适配
TOLA: Text-aware One-Step Latent Adaptation for Diffusion-based Text Image Super-Resolution
- Shanghai Jiao Tong University(上海交通大学)
- University of Chinese Academy of Sciences(中国科学院大学)
- Shanghai AI Laboratory(上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
TOLA提出一种文本感知单步潜在适配框架,通过置信度加权文本条件模块和轻量级潜在残差校正模块,在无需迭代图像-文本扩散的情况下,实现文本图像超分辨率,并在CTR-TSR-Test和RealCE-200基准上取得最优性能,PSNR至少提升2.72 dB。
AI中文摘要:
文本图像超分辨率(TSR)旨在未知退化条件下恢复视觉上忠实且可读的文本。现有的基于扩散的方法通常依赖于对高分辨率图像或其文本先验的多步预测,导致计算成本和推理延迟过高。更关键的是,错误的文本先验可能被反复注入去噪过程,使得图像和文本预测相互强化,并逐步将早期识别错误放大为清晰但语义错误的字符。为解决这些局限,我们提出TOLA,一种无需迭代图像-文本扩散的文本感知单步潜在适配框架。TOLA由两个关键模块组成。首先,一个置信度加权的文本条件模块仅构建一次语义条件,并在不可靠的OCR预测污染图像重建之前将其抑制。其次,一个轻量级的潜在残差校正模块显式估计并校正结构化残差误差,以恢复缺失或失真的笔画细节。大量实验表明,我们在CTR-TSR-Test(×4)和RealCE-200基准上的所有评估指标上均达到最先进性能。值得注意的是,我们的TOLA在CTR-TSR-Test上的PSNR持续超越现有基于扩散的TSR方法至少2.72 dB。
英文摘要:
Text image super-resolution (TSR) aims to recover visually faithful and readable text under unknown degradations. Existing diffusion-based methods typically rely on multi-step prediction of either the high-resolution image or its text prior, resulting in prohibitive computational cost and inference latency. More critically, an erroneous text prior may be repeatedly injected into the denoising process, causing image and text predictions to reinforce each other and progressively amplify an early recognition error into a sharp yet semantically incorrect character. To address these limitations, we propose TOLA, a Text-aware One-step Latent Adaptation framework without iterative image-text diffusion. TOLA consists of two key modules. First, a confidence-weighted text conditioning module constructs the semantic condition only once and suppresses unreliable OCR predictions before they contaminate image reconstruction. Second, a lightweight latent residual correction module explicitly estimates and corrects the structured residual errors to recover missing or distorted stroke details. Extensive experiments demonstrate our state-of-the-art performance across all evaluation metrics on both CTR-TSR-Test ($\times 4$) and RealCE-200 benchmarks. It is worth noting that our TOLA consistently surpasses existing diffusion-based TSR methods by at least 2.72 dB in PSNR on CTR-TSR-Test.