PaCo-RL:通过成对奖励建模推进强化学习用于一致图像生成
PaCo-RL: Advancing Reinforcement Learning for Consistent Image Generation with Pairwise Reward Modeling
浏览论文内容
中文总结 AI 辅助
本文提出PaCo-RL框架,结合专门的 consistency 奖励模型与高效强化学习算法,提升图像生成中的一致性与稳定性,实验表明其在视觉一致性上优于现有方法。
中文摘要 AI 辅助
一致的图像生成需要在多个图像中忠实保持身份、风格和逻辑一致性,这对故事讲述和角色设计等应用至关重要。监督训练方法由于缺乏大规模捕捉视觉一致性的数据集和建模人类感知偏好的复杂性而难以完成此任务。本文认为强化学习(RL)提供了一种有前途的替代方案,使模型能够在无数据的情况下学习复杂的主观视觉标准。为此,我们引入PaCo-RL,一个综合框架,结合专门的consistency奖励模型与高效的RL算法。第一个组件,PaCo-Reward,是通过自动化子图配对构建的大型数据集训练的成对一致性评估器,通过生成性、自回归评分机制增强任务意识指令和CoT原因进行一致性评估。第二个组件,PaCo-GRPO,利用一种新的分辨率解耦优化策略显著降低RL成本,同时利用日志驯服的多奖励聚合机制确保平衡和稳定的奖励优化。在两个代表性子任务上的广泛实验表明,PaCo-Reward显著提高了与人类视觉一致性感知的对齐程度,而PaCo-GRPO实现了最先进的视觉一致性性能,同时提高了训练效率和稳定性。这些结果突显了PaCo-RL作为一致图像生成的实用且可扩展的解决方案的潜力。项目页面可在https://x-gengroup.github.io/HomePage_PaCo-RL/上找到。
英文摘要
Consistent image generation requires faithfully preserving identities, styles, and logical coherence across multiple images, which is essential for applications such as storytelling and character design. Supervised training approaches struggle with this task due to the lack of large-scale datasets capturing visual consistency and the complexity of modeling human perceptual preferences. In this paper, we argue that reinforcement learning (RL) offers a promising alternative by enabling models to learn complex and subjective visual criteria in a data-free manner. To achieve this, we introduce PaCo-RL, a comprehensive framework that combines a specialized consistency reward model with an efficient RL algorithm. The first component, PaCo-Reward, is a pairwise consistency evaluator trained on a large-scale dataset constructed via automated sub-figure pairing. It evaluates consistency through a generative, autoregressive scoring mechanism enhanced by task-aware instructions and CoT reasons. The second component, PaCo-GRPO, leverages a novel resolution-decoupled optimization strategy to substantially reduce RL cost, alongside a log-tamed multi-reward aggregation mechanism that ensures balanced and stable reward optimization. Extensive experiments across the two representative subtasks show that PaCo-Reward significantly improves alignment with human perceptions of visual consistency, and PaCo-GRPO achieves state-of-the-art consistency performance with improved training efficiency and stability. Together, these results highlight the promise of PaCo-RL as a practical and scalable solution for consistent image generation. The project page is available at https://x-gengroup.github.io/HomePage_PaCo-RL/.