SPARK:用于基于冻结DiT的超分辨率的输入条件稀疏激活调制
SPARK: Input-Conditioned Sparse Activation Modulation for Frozen DiT-based Super-Resolution
浏览论文内容
中文总结 AI 辅助
该研究提出SPARK,通过轻量输入条件控制器仅调制冻结DiT超分辨率模型的少量主导通道,在三个基准数据集上实现保真度与感知质量的一致提升。
中文摘要 AI 辅助
现实世界的图像超分辨率(SR)越来越依赖扩散Transformer(DiT)作为骨干网络,其内部激活往往由少量规模巨大的通道主导。然而,要提升这些模型的感知质量,通常仍需要对网络进行微调或附加额外的适配器,这使得这种结构化的激活空间在适配过程中基本未被探索。我们研究主导通道是否可作为基于冻结DiT的SR模型的紧凑适配接口。我们首先表征预训练SR骨干中这些通道的行为,并通过可控干预证明它们对重建质量有强烈影响。基于此观察,我们提出SPARK,这是一种轻量的输入条件控制器,仅为选定通道预测有界的逐通道仿射变换,同时保持SR骨干和VAE冻结。主导通道通过在线激活排序过程识别,仅优化一个以低分辨率VAE潜变量为条件的小型预测器。在三个基于DiT的SR骨干网络上,针对DIV2K、RealSR和DRealSR数据集的实验表明,在每个流和每个块仅调制8个通道的情况下,SPARK在保真度和感知质量方面均取得一致提升。可控对比进一步表明,这些提升无法仅用参数预算或选定通道的访问来解释。
英文摘要
Real-world image super-resolution (SR) increasingly relies on Diffusion Transformer (DiT) backbones, whose internal activations can be dominated by a small number of massive channels. Yet improving perceptual quality in these models still typically requires fine-tuning the network or attaching additional adapters, leaving this structured activation space largely unexplored for adaptation. We investigate whether dominant channels can instead serve as a compact adaptation interface for frozen DiT-based SR models. We first characterize their behavior in pretrained SR backbones and show through controlled interventions that they strongly affect reconstruction quality. Building on this observation, we introduce SPARK, a lightweight input-conditioned controller that predicts bounded per-channel affine transformations for only the selected channels, while keeping the SR backbone and VAE frozen. Dominant channels are identified through an online activation-ranking procedure, and only a small predictor conditioned on the low-resolution VAE latent is optimized. Experiments on three DiT-based SR backbones across DIV2K, RealSR, and DRealSR show consistent gains in both fidelity and perceptual quality while modulating only eight channels per stream and block. Controlled comparisons further show that these gains cannot be explained by parameter budget or access to the selected channels alone.
发表机构
- University of Modena and Reggio Emilia(摩德纳和雷焦艾米利亚大学)
机构由 AI 辅助整理,请以论文原文为准。