Policy-based Tuning of Autoregressive Image Models with Instance- and Distribution-Level Rewards
基于策略的自回归图像模型实例和分布级奖励调优
机构 * Middle East Technical University (METU)(中东部技术大学) ; University of Southern Denmark(南部丹麦大学)
AI总结 本文提出一种轻量级强化学习框架,通过将基于标记的自回归合成视为马尔可夫决策过程,结合分布级LOO-FID奖励和实例级CLIP/HPSv2奖励,提升图像生成的质量和多样性。
Comments Accepted at the European Conference on Computer Vision (ECCV), 2026. This version includes the appendix