发表机构
The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究文本到图像生成中样本模式覆盖不足问题,提出多轴最大@K强化学习目标,通过特定信用分配机制改善覆盖,在感知外观公平性评估中提高公平分数,且保持图像质量和文本对齐。
AI 中文摘要
文本到图像(T2I)模型能合成逼真且符合提示的图像,但同一提示生成的样本往往只涵盖视觉上不同模式的一小部分。这限制了图像多样性,对于以人物为中心的提示,可能反映或放大人口统计学偏差。我们将此问题形式化为预定义语义指定模式集的覆盖,即目标模式覆盖。然后提出多轴最大@K,一种基于组的强化学习目标,用于改善基于扩散的T2I模型中的这种覆盖。给定一组样本和每个目标类别的一个分数,多轴最大@K首先对每个类别取样本中的最大分数,然后将这些类别最大值相加。结果信用分配仅在样本增加该类别的组内最大值时给予其正权重,允许不同样本对不同类别做出贡献。我们首先在合成混合物和SD3.5 - M上使用基于确定性像素的颜色奖励验证信用分配机制。然后在感知外观公平性上评估相同目标。在三个针对保留提示的自动评估器上,多轴最大@K相对于基础模型将公平分数提高了0.23 - 0.36,同时保持图像质量和文本对齐。
英文摘要
Text-to-image (T2I) models can synthesize realistic, prompt-aligned images, yet samples generated for the same prompt often cover only a small subset of visually distinct modes. This limits diversity and, for person-centric prompts, can reflect or amplify demographic skew. We formalize this problem as target-mode coverage, the coverage of a predefined set of semantically specified modes, and propose multi-axis max@K, a group-based reinforcement learning objective for improving it in diffusion-based T2I models. Given a group of samples and one score per target mode, multi-axis max@K first takes the maximum score across samples for each mode and then sums these per-mode maxima. The resulting credit assignment gives a sample positive weight on a mode only when it raises that mode's group maximum, so different samples can contribute to different modes. We validate the credit-assignment mechanism on a synthetic mixture and on SD3.5-M with deterministic pixel-based color rewards, and then apply the same objective to perceived-appearance fairness. On held-out prompts, multi-axis max@K improves the Fairness Score by 0.23-0.36 over the base model under three automatic evaluators, while maintaining image quality and text alignment. Code is available at https://github.com/KuOnoda/multi-axis-maxk.
CommentsAccepted at WACV 2027