arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05109cs.CV

LoopMoEVR:基于循环的退化感知混合专家模型用于统一超高清视频恢复

LoopMoEVR: Loop-Based Degradation-Aware Mixture-of-Experts for Unified UHD Video Restoration

  • Shandong Normal University(山东师范大学)
  • China Academy of Information and Communications Technology(中国信息通信研究院)
  • Central South University(中南大学)
  • Nanjing University of Science and Technology(南京理工大学)
  • Hebei University of Technology(河北工业大学)
  • Shandong University of Finance and Economics(山东财经大学)
  • National University of Defense Technology(国防科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Yucheng Xin, Runci Bai, Yongcong Wang, Guangwei Gao, Jiao Liu, Dianjie Lu, Linwei Fan, Zhuoran Zheng

AI总结:

针对现有高清恢复模型深度过大而收益有限的问题,提出基于循环的混合专家模型LoopMoEVR,以约0.884M参数统一处理超高清视频去雾、去雨、去噪和低光增强,达到最先进性能。

AI中文摘要:

近年来,统一高清图像恢复引起了广泛关注;然而,现有模型为了追求更好的泛化能力,往往过度增加网络深度,但收益却有限。与此同时,基于循环的学习范式因其低参数数量和强大的回归能力而受到广泛关注,例如GPT-6和循环Transformer。在本文中,我们将循环学习范式引入到需要跨领域学习的恢复任务中。具体而言,我们提出了LoopMoEVR,一种基于循环的混合专家模型,能够处理退化的超高清(UHD)输入。首先,我们设计了一种退化条件化的低秩循环嵌入,用于构建依赖输入的分阶段条件。其次,开发了一种时空迭代自适应归一化模块(称为IterAda3DN),用于融合局部特征与全局循环上下文,从而执行逐位置仿射调制。最后,专家分支进一步将注意力更新的局部和全局视频状态与循环条件相结合,生成专用的调制参数,同时,一个输入条件化的深度预测器自适应地配置循环迭代次数。该模型仅需约0.884M可训练参数,即可统一处理超高清视频去雾、去雨、去噪和低光增强任务,在公共基准和真实场景中均达到了最先进的恢复性能。

英文摘要:

Recently, unified high-definition image restoration has attracted considerable attention; however, existing models tend to excessively increase their depth in pursuit of improved generalization, which often yields only limited gains. Meanwhile, loop-based learning paradigms have drawn widespread attention due to their low parameter counts and strong regression capability, as exemplified by GPT-6 and looped Transformers. In this paper, we introduce the loop learning paradigm to address restoration tasks that require cross-domain learning. Specifically, we propose LoopMoEVR, a loop-based mixture-of-experts model capable of handling degraded ultra-high-definition (UHD) inputs. First, a degradation-conditioned low-rank loop embedding is designed to construct input-dependent stage conditions. Second, a spatio-temporal iterative adaptive normalization module, termed IterAda3DN, is developed to fuse local features with global loop context, thereby performing position-wise affine modulation. Finally, the expert branches further integrate the attention-updated local and global video states with the loop conditions to generate dedicated modulation parameters, while an input-conditioned depth predictor adaptively configures the number of loop iterations. With only approximately 0.884M trainable parameters, the proposed model uniformly handles UHD video dehazing, deraining, denoising, and low-light enhancement tasks, achieving state-of-the-art restoration performance on both public benchmarks and real-world scenarios.

↑