发表机构
Huawei; Graz University of Technology(华为; 格拉茨技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出表示特征适配器(RFA),通过调整DINO特征中的残差来去除域差异,仅训练290万参数的适配器,以极低的训练成本在图像转换任务上达到或超越现有方法。
AI 中文摘要
清除视频中的雾、雨或雪,或将渲染图转换为照片,必须去除源域并保留场景。非配对转换器通过其生成器看到源外观(像素、近可逆的潜在表示或控制图)并保留它来实现这一点。DINO特征图固定了场景中的内容,并将天气、光照和渲染风格作为特征范数的13%至14%的残差携带。我们提出了表示特征适配器(RFA),一个拥有290万参数的网络,用于移动这一残差。我们仅训练适配器及其判别器;编码器和特征条件解码器,针对所有条件训练一次,保持冻结。与CycleGAN-Turbo相比,它在雾和夜间场景的KID指标上均领先,在雪、雨和雾霾场景中与噪声水平相当。在模拟到真实转换中,它在两个指标上均领先REGEN和HyPER-GAN。只有RFA能在保持场景的同时去除雨水。去除过程会损失场景结构:CycleGAN-Turbo在除雾以外的所有条件下保留更多结构。在VAE潜在表示上,相同的适配器退化为恒等映射,而来自其他组从未见过其输出的解码器能渲染其输出。RFA的可训练参数比CycleGAN-Turbo少约160倍,且每种条件的训练时间不足其五分之一。
英文摘要
Clearing fog, rain or snow from footage, or turning renders into photographs, must remove the source domain and keep the scene. Unpaired translators carry it through because their generator sees the source appearance (pixels, a near-invertible latent or a control map) and keeps it. A DINO feature map fixes what is in the scene and carries weather, lighting and rendering style as a residue of 13 to 14% of the feature norm. We propose the Representation Feature Adapter (RFA), a 2.9M-parameter network that moves this residue. We train only the adapter and its discriminators; the encoder and a feature-conditioned decoder, trained once for all conditions, stay frozen. Against CycleGAN-Turbo it is ahead on both metrics on fog and on KID on night, and level within noise on snow, rain and haze. On sim-to-real it leads REGEN and HyPER-GAN on both metrics. Only the RFA removes the rain while keeping the scene. The removal costs scene structure: CycleGAN-Turbo keeps more on every condition but fog. On VAE latents the identical adapter collapses to the identity, and decoders from other groups that never saw it render its output. The RFA has about 160 times fewer trainable parameters than CycleGAN-Turbo and under a fifth of its per-condition training time.
Comments9 pages main text, 28 pages including appendix. 12 figures, 13 tables