LucidNFT: 基于LR锚定的多奖励偏好优化用于基于流的现实世界超分辨率
LucidNFT: LR-Anchored Multi-Reward Preference Optimization for Flow-Based Real-World Super-Resolution
浏览论文内容
中文总结 AI 辅助
本文提出LucidNFT框架,通过引入LucidConsistency、解耦奖励归一化策略和LucidLR数据集,解决现实世界超分辨率中LR参考一致性、奖励优化瓶颈和真实退化覆盖不足的问题,提升感知质量。
中文摘要 AI 辅助
生成现实世界图像超分辨率(Real-ISR)可以从严重退化的低分辨率(LR)输入中合成视觉逼真的细节,但其随机采样导致关键失败模式难以避免:输出可能看起来清晰但不忠实于LR证据,出现语义或结构幻觉。基于偏好强化学习(RL)是自然选择,因为每个LR输入会产生一组候选修复。然而,Real-ISR的有效对齐受到三个耦合挑战的阻碍:(i)缺乏一个鲁棒于退化但对局部幻觉敏感的LR参考忠实信号;(ii)rollout组优化瓶颈,其中在归一化前标量化异质奖励会压缩目标对比度并削弱扩散NFT风格的奖励加权更新;(iii)真实退化的有限覆盖,限制了rollout多样性和偏好信号质量。我们提出LucidNFT,一种用于流匹配Real-ISR的多奖励RL框架。LucidNFT引入LucidConsistency,一种退化不变且幻觉敏感的LR参考评估器,通过内容一致的退化池和原始输入修补的硬负样本训练;一种解耦的奖励归一化策略,该策略在融合前在每个LR条件化的rollout组内保留目标对比度;以及LucidLR,一个大规模的真实世界退化图像集合用于鲁棒的RL微调。大量实验表明,LucidNFT在强大的基于流的Real-ISR基线中提升了感知质量,同时在多样化的现实世界场景中普遍保持LR参考一致性。
英文摘要
Generative real-world image super-resolution (Real-ISR) can synthesize visually convincing details from severely degraded low-resolution (LR) inputs, yet its stochastic sampling makes a critical failure mode hard to avoid: outputs may look sharp but be unfaithful to the LR evidence, exhibiting semantic or structural hallucinations. Preference-based reinforcement learning (RL) is a natural fit because each LR input yields a rollout group of candidate restorations. However, effective alignment in Real-ISR is hindered by three coupled challenges: (i) the lack of an LR-referenced faithfulness signal that is robust to degradation yet sensitive to localized hallucinations, (ii) a rollout-group optimization bottleneck where scalarizing heterogeneous rewards before normalization compresses objective-wise contrasts and weakens DiffusionNFT-style reward-weighted updates, and (iii) limited coverage of real degradations, which restricts rollout diversity and preference signal quality. We propose LucidNFT, a multi-reward RL framework for flow-matching Real-ISR. LucidNFT introduces LucidConsistency, a degradation-invariant and hallucination-sensitive LR-referenced evaluator trained with content-consistent degradation pools and original-inpainted hard negatives; a decoupled reward normalization strategy that preserves objective-wise contrasts within each LR-conditioned rollout group before fusion; and LucidLR, a large-scale collection of real-world degraded images for robust RL fine-tuning. Extensive experiments show that LucidNFT improves perceptual quality on strong flow-based Real-ISR baselines while generally maintaining LR-referenced consistency across diverse real-world scenarios.
发表机构
- The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
- The Hong Kong University of Science and Technology(香港科学与技术大学)
机构由 AI 辅助整理,请以论文原文为准。