Noise2Noise 再探:训练配对分布主导自监督去噪中的损失函数选择
Noise2Noise Revisited: Training Pair Distributions Dominate Loss Choice in Self-Supervised Denoising
AI总结:
本文通过实验证明,在自监督去噪中,训练配对分布而非损失函数选择(L1 vs L2)主导性能,并给出 Kodak24 与 SIDD 上的定量证据。
AI中文摘要:
Noise2Noise (N2N) 在独立损坏的观测值对上训练去噪器,从而消除了对干净参考图像的需求。我们压力测试了关于 L1 损失在此优于 L2 损失的两个自然猜想。首先,关于 L1 损失通过参数稀疏性赋予鲁棒性的假设,混淆了损失与 Lasso 正则化:显式的 Lasso 惩罚产生了预期的稀疏性,但未能重现 L1 的跨噪声行为,而 L1 和 L2 训练得到的权重分布无法区分。其次,对于对称的信号后验分布,两种损失的总体最优值完全一致,对于集中的后验分布则几乎一致。因此,测得的差异主要由优化动态(有界影响梯度)主导,我们通过梯度统计和污染目标训练对此进行了探究。在 Kodak24 数据集上,使用五种合成噪声族,L1 损失相对于 L2 具有统计显著的微弱优势,低于 1 dB PSNR,且在 14 个噪声列中的 13 个上,三种随机种子下均保持一致。对于真实相机噪声,损失并非分布中的决定性变量:在官方 SIDD 验证块上,无论使用何种损失,合成高斯训练的 N2N 模型仅比噪声输入提升 0.8 至 3.7 dB,而在 SIDD 自身的噪声对上重新训练(从不读取真实值)则提升 9.4 至 11.0 dB,远超 BM3D。所有指标均基于原始网络输出,本研究不声称排行榜排名。训练配对分布而非损失函数承载了归纳偏置。这一设计规则适用于任何无法获得干净参考图像的场景,从显微镜到工业检测传感器。
英文摘要:
Noise2Noise (N2N) trains denoisers on pairs of independently corrupted observations, eliminating clean references. We stress-test two natural conjectures about why the L1 loss outperforms L2 here. First, the hypothesis that the L1 loss confers robustness via parameter sparsity confuses the loss with Lasso regularization: an explicit Lasso penalty produces the predicted sparsity yet fails to reproduce L1's cross-noise behavior, while L1- and L2-trained weight distributions are indistinguishable. Second, the population optima of the two losses coincide exactly for symmetric signal posteriors and nearly so for concentrated ones. Measured differences are therefore dominated by optimization dynamics (bounded-influence gradients), which we probe with gradient statistics and contaminated-target training. On Kodak24 with five synthetic noise families, the L1 loss holds a statistically significant edge over L2, below 1 dB PSNR, holding across three seeds on 13 of the 14 noise columns. On real camera noise the loss is not the decisive variable in distribution: on official SIDD validation blocks, synthetic-Gaussian-trained N2N models gain only 0.8 to 3.7 dB over the noisy input regardless of loss, while retraining on SIDD's own noisy pairs, never reading ground truth, gains 9.4 to 11.0 dB, far ahead of BM3D. All metrics are on raw network outputs, and the study makes no leaderboard claim. The training pair distribution, not the loss, carries the inductive bias. That design rule applies wherever clean references are unobtainable, from microscopy to industrial inspection sensors.