arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28730eess.IVcs.CV

基于均值流的小波域高效JPEG恢复方法

Efficient JPEG Restoration in the Wavelet Domain via Mean Flows

Stefan-Alexandru Asandei, Mihai-Alexandru Radu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出含65M参数的生成式JPEG恢复器,以两级Haar变换和秩增强线性注意力DiT为核心,在保障低LPIPS的同时实现高吞吐量,适用于设备端部署。

中文摘要 AI 辅助

现有JPEG恢复系统采用大模型虽能实现较高质量,但通常速度过慢、成本过高,难以高效部署于设备端。本文提出一种含65M参数的生成式恢复器,在LIVE-1、Urban100和DIV2K验证集上,当量化因子(QF)为10和20时,取得最低的LPIPS值;同时在单张RTX 3090显卡上,处理1024×1024分辨率图像时保持8.05张/秒的吞吐量,约为单步SODiff吞吐量的4.9倍,而参数仅为其1/20。该模型从零开始训练,用完全可逆的两级Haar变换替代学习型VAE编解码器,通过秩增强线性注意力DiT(内部估算压缩严重程度)预测干净的小波域残差,并采用改进的MeanFlow目标函数优化,可在1或2次网络评估中完成推理,无需蒸馏。在严重压缩(QF 5)下,大型预训练先验仍表现更强,而本文模型则优先保障受部署限制的恢复任务的吞吐量。

英文摘要

Latest JPEG restoration systems achieve strong quality with large models, yet often remain too slow and expensive for efficient on-device deployment. We present a 65M-parameter generative restorer that attains the lowest LPIPS at QF 10 and 20 on LIVE-1, Urban100, and DIV2K-val while sustaining 8.05 images/s at $1024\times1024$ on a single RTX 3090, roughly $4.9\times$ the reported throughput of one-step SODiff at one-twentieth of its parameters. Trained from scratch, the model replaces the learned VAE encoder-decoder with an exactly invertible two-level Haar transform, predicts a clean wavelet-domain residual through a rank-enhanced linear-attention DiT that estimates compression severity internally, and is optimized with an improved MeanFlow objective that enables inference in one or two network evaluations without distillation. Large pretrained priors remain stronger under severe compression (QF 5), whereas our model prioritizes throughput for deployment-constrained restoration.

发表机构

  • Faculty of Physics, "Alexandru Ioan Cuza" University of Iaşi(雅西亚历山德鲁·约安·库扎大学物理学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑