PixelSR:基于像素分类的高效屏幕内容超分辨率方法
PixelSR: Efficient Screen Content Super-Resolution via Pixel Classification
浏览论文内容
中文总结 AI 辅助
PixelSR通过训练阶段的像素分箱分类计算内容注意力、测试阶段利用屏幕内容特性优化推理,在屏幕内容超分辨率任务中实现了性能先进且推理速度更快的效果。
中文摘要 AI 辅助
屏幕内容图像通常由文本和图形构成,与自然图像相比,这类人造图像包含大量清晰且重复的结构。然而,现有的屏幕内容超分辨率研究未能充分利用屏幕内容的特殊特性,在模型性能和推理速度上仍有很大提升空间。本文提出PixelSR这一简单却有效的方法,以提升超分辨率性能并加快推理速度。为提升模型性能,我们在训练阶段通过像素分箱对像素进行分类,以计算内容注意力。具体而言,将像素分箱为内容相关组后,从每组内的像素特征中聚合内容注意力,为每个像素引入内容相关的非局部感受野。在测试阶段,我们利用屏幕内容的自重复性和冗余性,在不损失模型性能的前提下加快推理速度。针对每张测试图像,我们将目标高分辨率像素分为三类:唯一像素、重复像素和背景像素。对唯一像素进行常规网络处理,并将其预测结果缓存至动态查找表;对于已在唯一像素中出现过的重复像素,直接从查找表中检索预测结果,无需网络处理;对于背景像素,采用最近邻算法生成高分辨率像素。动态查找表会被清空,再对下一张测试图像重复上述流程。实验表明,我们的PixelSR在屏幕内容超分辨率任务中达到了先进性能,且推理时间更短。
英文摘要
Screen content images are generally composed of texts and graphics. Compared to natural images, these man-made images contain a large quantity of sharp but repetitive structures. However, existing works in screen content super-resolution underutilize the special characteristics of screen content, leaving a large room to improve model performance and speed up. In this paper, we propose PixelSR, a simple yet effective method to improve super-resolution performance but with faster inference speed. To improve model performance, we classify pixels via pixel binning to compute content attention in the training phase. Specifically, after binning pixels into content-dependent groups, content attention is aggregated from pixel features within each group to introduce a content-dependent and non-local receptive field for every pixel. In the testing phase, we utilize the properties of self-repetitiveness and redundancy in screen content to speed up inference without the loss of model performance. We divide targeted high-resolution pixels into three types, which are unique pixels, repeated pixels, and background pixels for each test image. We conduct conventional network processing on unique pixels and cache their predictions in the on-the-fly lookup table. For repeated pixels which have appeared in unique pixels, we directly retrieve prediction results from the lookup table without network processing. For background pixels, we use the nearest neighbor algorithm to generate high-resolution pixels. The on-the-fly lookup table is cleaned and repeats the procedure above for the next test image. Experiments show our PixelSR achieves state-of-the-art performance with shorter inference time in screen content super-resolution.
发表机构
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。