arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02967cs.CVcs.AI

通过组合偏好与规则奖励对前沿文生图模型进行后训练

Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards

发表机构竞技场智能公司 · 加州大学洛杉矶分校
查看机构详情
  • Arena Intelligence Inc(竞技场智能公司)
  • UCLA(加州大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

Yuanhao Ban, I-Hung Hsu, Anastasios Angelopoulos, Wei-Lin Chiang, Ion Stoica, Cho-Jui Hsieh

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种组合偏好与规则奖励的后训练方案,用于提升文生图模型性能,在Arena排行榜上显著超越基线,并发布训练数据子集以促进可复现研究。

中文摘要 AI 辅助

最近的文生图模型已经取得了显著的视觉质量,但通过后训练来改进它们仍然具有挑战性,因为没有任何单一的奖励信号能够捕捉人类偏好的全部范围。在这项工作中,我们为开放域文生图开发了一种简单有效的后训练方案,该方案基于互补奖励信号的组合。我们的奖励系统由两个主要部分组成:一个偏好奖励,使用Bradley-Terry目标在大规模人类偏好数据上进行训练,以捕捉整体的人类美学和感知偏好;以及基于规则的奖励,这些奖励明确评估提示忠实度和其他理想属性,同时提供针对奖励黑客的防护措施。一个关键挑战是如何组合这些异构的奖励信号。我们表明,简单的加权平均会导致次优的优化行为,并提出了一种简单的奖励组合策略,该策略能更有效地平衡偏好优化与规则满足。在Arena文生图排行榜(此 https URL )上,我们的RL训练的Flux2dev的Elo评分比基础模型高出69分,而我们后训练的Ideogram-4超过了排行榜上的所有开源模型,达到了1223.5的Elo评分。(关于最先进性能的声明基于截至2026年9月4日的Arena排行榜快照。)我们的结果表明,有效的奖励用于前沿生成模型训练需要广泛覆盖用户意图,并且对优化下的利用具有鲁棒性。为了支持可重现的研究,我们发布了Arena-T2I-Training,一个包含1K子集的训练数据,该子集恢复了全规模训练的部分收益,提供了我们希望将促进未来文生图模型后训练工作的资源。

英文摘要

Recent text-to-image generation models have achieved remarkable visual quality, but improving them through post-training remains challenging because no single reward signal captures the full range of human preference. In this work, we develop a simple and effective post-training recipe for open-domain text-to-image generation based on the composition of complementary reward signals. Our reward system consists of two main components: a preference reward, trained on large-scale human preference data using a Bradley-Terry objective to capture overall human aesthetic and perceptual preferences, and rubric-based rewards, which explicitly evaluate prompt faithfulness and other desirable properties while providing safeguards against reward hacking. A key challenge is how to combine these heterogeneous reward signals. We show that a naive weighted average leads to suboptimal optimization behavior, and propose a simple reward composition strategy that more effectively balances preference optimization with rubric satisfaction. In the Arena text-to-image leaderboard (https://arena.ai/), our RL-trained Flux2dev achieves an Elo rating 69 points above the base model, and our post-trained Ideogram-4 surpasses every open-source model on the leaderboard, reaching an Elo of 1223.5. (Claims of state-of-the-art performance are based on the Arena leaderboard snapshot as of September 4, 2026.) Our results suggest that effective rewards for frontier generative-model training require broad coverage of user intent and robustness to exploitation under optimization. To support reproducible research, we release Arena-T2I-Training, a 1K subset of training data that recovers some gains of full-scale training, providing a resource that we hope will facilitate future work on post-training for text-to-image models.

↑