发表机构
University of Central Florida; Amazon(中佛罗里达大学; 亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SatisDive,一种无需训练的推理时方法,通过设置奖励下限和多样性阈值,在文本到图像扩散中遍历满意度-多样性帕累托前沿,显著提升最差候选奖励。
AI 中文摘要
文本到图像生成使用户能够探索由同一提示生成的多个图像。为了使这些生成的图像有用,每个图像都必须反映用户的偏好(通过学习的奖励衡量),并在视觉上与其他图像不同以保持多样性。现有方法存在局限性:它们要么分别处理奖励和多样性,要么将两者合并为一个总分,使得高多样性可以抵消低奖励。在本文中,我们通过将生成表述为满意(satisficing)来解决这些局限性:每个图像(候选)必须满足奖励下限,并且图像批次必须满足多样性阈值。奖励下限控制最差候选奖励与批次多样性之间的平衡;我们证明改变该下限定义了一个帕累托前沿。为了遍历这一前沿,我们引入了SatisDive,一种无需训练、推理时的方法。SatisDive使用批次相对奖励截止值来区分较低和较高奖励的候选,强调对低于截止值的候选进行奖励改进,并对高于截止值的候选强调多样性。在Pick-a-Pic上,在匹配的DreamSim下,以FLUX.1-dev为基础模型和HPSv3为奖励,SatisDive将最差候选奖励相对于FK steering提高了最多0.43;以SANA-1.6B为基础模型和ImageReward为奖励,提高了最多0.70。更广泛地说,在它们重叠的DreamSim范围内,SatisDive的满意度-多样性曲线在每个设置中都帕累托支配FK steering的曲线。
英文摘要
Text-to-image generation enables users to explore several images generated from the same prompt. For these generated images to be useful, each one must reflect the user's preferences, measured by a learned reward, and differ visually from the others to maintain diversity. Existing methods are limited: they either address reward and diversity separately or combine them in one aggregate score, enabling high diversity to offset low rewards. In this paper, we address these limitations by formulating generation as satisficing: every image (candidate) must satisfy a reward floor and the batch of images must satisfy a diversity cutoff. The reward floor controls the balance between worst-candidate reward and batch diversity; we show that varying this floor defines a Pareto frontier. To traverse this frontier, we introduce SatisDive, a training-free inference-time method. SatisDive uses a batch-relative reward cutoff to distinguish lower- from higher-reward candidates, emphasizing reward improvement for candidates below the cutoff and diversity among candidates above it. On Pick-a-Pic, at matched DreamSim, SatisDive improves worst-candidate reward over FK steering by up to 0.43 with FLUX.1-dev as the base model and HPSv3 as the reward, and by up to 0.70 with SANA-1.6B as the base model and ImageReward as the reward. More broadly, across their overlapping DreamSim ranges, SatisDive's satisfaction-diversity curve Pareto-dominates FK steering's curve in each setting.
Comments40 pages, including appendices. Code: https://github.com/UCF-CRCV/SatisDive