用于屏幕内容视频质量增强的时空多尺度网络
Spatial-Temporal Multi-scale Network for Screen Content Video Quality Enhancement
另 2 家 · 查看机构详情
- Shenzhen Polytechnic University(深圳职业技术大学)
- The Hong Kong Polytechnic University(香港理工大学)
- Hong Kong Chu Hai College(香港珠海学院)
- Dongguan University of Technology(东莞理工学院)
- City University of Hong Kong(香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对屏幕内容视频因突然运动和高频细节导致传统增强方法失效的问题,提出时空多尺度网络(STM-Net),通过三个互补模块(PG-STD、BTFE、CMFD)实现增强,在客观和主观质量上超越现有方法。
中文摘要 AI 辅助
与自然视频不同,屏幕内容视频(SCVs)具有突然运动、场景切换以及文本和图形等高频率细节的特点。传统的视频增强方法严重依赖时间连续性,在处理SCVs时往往因时间相关性被破坏而导致性能下降。为解决这些挑战,我们提出了时空多尺度网络(STM-Net),一种专门针对压缩SCV增强的新框架。我们的方法整合了三个互补组件:先验引导的时空调度器(PG-STD)将输入路由到三个并行流以避免特征污染,双向时间特征提取(BTFE)模块自适应处理突然转换而无需显式检测,以及级联多尺度特征蒸馏(CMFD)模块保留关键的高频细节。实验结果表明,STM-Net在客观指标和主观视觉质量上均优于现有最先进方法,为屏幕内容伪影提供了稳健的解决方案。代码可在该https URL获取。
英文摘要
Different from natural videos, Screen Content Videos (SCVs) are characterized by abrupt motion, scene switches, and high-frequency details such as text and graphics. Conventional video enhancement methods, which rely heavily on temporal continuity, often suffer from performance degradation when processing SCVs due to the disruption of temporal correlations. To address these challenges, we propose the Spatial-Temporal Multi-scale Network (STM-Net), a novel framework specifically tailored for compressed SCV enhancement. Our approach integrates three complementary components: a Prior-Guided Spatio-Temporal Dispatcher (PG-STD) that routes input into three parallel streams to avoid feature contamination, a Bidirectional Temporal Feature Extraction (BTFE) module that adaptively handles abrupt transitions without explicit detection, and a Cascaded Multi-scale Feature Distillation (CMFD) module that preserves critical high-frequency details. Experimental results demonstrate that STM-Net outperforms state-of-the-art methods in both objective metrics and subjective visual quality, providing a robust solution for screen content artifacts. Code is available at https://github.com/HUANGZiyin1/STM-Net.