arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向场景文本图像超分辨率的耦合连续-离散生成方法

Coupled Continuous-Discrete Generation for Scene Text Image Super-Resolution

Axi Niu, Knag Zhang, Qingsen Yan, Hao Jin, Jinqiu Sun, Yanning Zhang

arXiv 2608.04525首次发表:更新:

AI 中文总结

针对现有场景文本图像超分辨率(STISR)方法误差传播、成本高的问题,提出统一框架DualTSR,通过耦合连续-离散生成实现高效准确的STISR,在多数据集上表现优于对比方法。

AI 中文摘要

场景文本图像超分辨率(STISR)旨在从退化输入中恢复视觉上合理的外观,同时保留字符语义。现有STISR系统通常依赖外部生成的先验或独立的图像与文本模型,导致误差传播和多阶段推理成本高昂。我们提出DualTSR,一种将STISR建模为耦合连续-离散生成的统一框架:条件流匹配恢复连续图像潜变量,吸收态离散扩散重构文本 token,两个过程共享一个多模态Transformer骨干,使图像与文本状态在整个生成过程中交互,推理时无需外部OCR先验。在CTR-TSR数据集上,DualTSR在X2和X4缩放倍数下均实现了对比方法中最佳的FID、LPIPS、ACC和NED;在对齐的RealCE子集上,它取得了最佳的FID、ACC和NED,且LPIPS具有竞争力。与X4缩放下的DiffTSR相比,DualTSR将ACC提升了12.78个百分点,同时将参数数量从12.3亿减少至2.03亿,端到端延迟从13.3秒降至132毫秒。这些结果表明DualTSR是一种准确且高效的STISR方法。

英文摘要

Scene text image super-resolution (STISR) aims to recover visually plausible appearance while preserving character semantics from degraded inputs. Existing STISR systems often rely on externally generated priors or separate image and text models, resulting in error propagation and costly multi-stage inference. We present DualTSR, a unified framework that formulates STISR as coupled continuous-discrete generation. Conditional flow matching restores continuous image latents, while absorbing-state discrete diffusion reconstructs text tokens. Both processes share a multimodal transformer backbone, allowing the evolving image and text states to interact throughout generation without an external OCR prior at inference. On CTR-TSR, DualTSR achieves the best FID, LPIPS, ACC, and NED among the compared methods at both X2 and X4. On an aligned RealCE subset, it obtains the best FID, ACC, and NED with competitive LPIPS. Compared with DiffTSR at X4, DualTSR improves ACC by 12.78 percentage points while reducing the parameter count from 1.23B to 203M and end-to-end latency from 13.3s to 132ms. These results establish DualTSR as an accurate and efficient method for STISR.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑