发表机构
Southern University of Science and Technology; Tencent Youtu Lab(南方科技大学; 腾讯优图实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ReNFT 是一种从扩散生成器内部修复奖励后训练模式崩溃的方法,在保留 NFT 高奖励的同时显著提升了 DreamSim-Div,为外部干预提供互补方案。
AI 中文摘要
扩散生成器的奖励后训练不可避免地会将概率质量集中在少数奖励偏好的模式上,这种模式崩溃会消除提示内多样性。现有的缓解崩溃的方法依赖外部信号或接口,用感知目标增强奖励、调整参考正则化或修改文本编码器,但 none 修复已崩溃的适配器同时保留获得的奖励。我们观察到在线后训练主要重新分配预训练继承的能力上的概率质量,而非学习新视觉内容,因此崩溃是抑制而非删除,可从生成器内部逆转。我们提出 ReNFT,通过内部概率质量重新校准修复高奖励、低多样性的适配器。无条件探测首先优先选择“反中心”提示,其中与提示无关的偏差最易暴露;两条策略主导的混合路线随后从同一提示和初始噪声生成匹配的反事实提议,一条探测冻结基础方向以寻找被抑制的替代方案,另一条暴露后训练的无条件倾向。带自适应翻转保护的奖励排名分配拉取和推送角色,联合配对的 NFT 更新实现修复。在 PickScore 和 GenEval 上,ReNFT 保留了 NFT 98.9%和 99.0%的奖励,同时 DreamSim-Div 分别提升了 58.8%和 55.0%,为外部干预提供了互补替代方案。
英文摘要
Reward post-training of diffusion generators inevitably concentrates probability mass on a few reward-favored modes, a mode collapse that erases within-prompt diversity. Existing methods for mitigating collapse rely on external signals or interfaces, augmenting the reward with perceptual objectives, adjusting reference regularization, or modifying the text encoder, but none repairs an adapter that has already collapsed while preserving the acquired reward. We observe that online post-training primarily reallocates probability mass over capabilities inherited from pretraining rather than learning new visual content. Collapse is therefore suppression, not deletion, and can be reversed from within the generator. We propose ReNFT, which repairs a high-reward, low-diversity adapter through internal probability-mass recalibration. Unconditional probes first prioritize "anti-hub" prompts where the prompt-independent bias is easiest to expose. Two policy-dominated mixed routes then generate matched counterfactual proposals from the same prompt and initial noise, one probing the frozen base direction for suppressed alternatives and the other exposing the post-trained unconditional tendency. Reward ranking with an adaptive flipping guard assigns pull and push roles, and a joint-and-paired NFT update realizes the repair. On PickScore and GenEval, ReNFT retains 98.9% and 99.0% of NFT's reward while improving DreamSim-Div by 58.8% and 55.0%, respectively, offering a complementary alternative to external interventions.
Comments17 pages, 13 figures, 4 tables. Project Page: https://yusenbao01.github.io/renft/