发表机构
Samsung Electronics; Seoul National University; Konkuk University(三星电子; 首尔大学; 建国大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种无需训练的扩散模型记忆缓解方法,通过高斯平滑重分配交叉注意力并强化内容令牌,在保持提示对齐的同时减少训练图像复现,并在多个模型上验证了帕累托最优性。
AI 中文摘要
文本到图像扩散模型在图像合成方面取得了显著进展,但可能通过紧密复现个别训练样本来表现出记忆现象。有效的缓解措施必须保留有用的提示信息以引导替代性描绘。我们提出了一种无需训练的方法,该方法在强化内容令牌贡献并减弱填充令牌贡献之前,通过高斯平滑重新分配交叉注意力,且无需额外的去噪器评估。通过这种干预,更强的内容条件可以在训练图像相似度相当的情况下改善提示对齐。局部分析确定了强化何时保留共享价值信息,而重新分配则减少局部注意力质量。在Stable Diffusion v1.4和v2.0上,所有评估的平滑宽度均位于训练图像相似度与提示对齐及图像偏好之间的经验帕累托前沿上。在Stable Diffusion上选择的配置无需进一步调整即可减少DeepFloyd IF中的模板复现。这些发现支持联合控制条件分配和强度以生成提示一致的替代方案。
英文摘要
Text-to-image diffusion models have achieved remarkable progress in image synthesis, yet can exhibit memorization by closely reproducing individual training examples. Effective mitigation must preserve useful prompt information to guide alternative depictions. We introduce a training-free method that redistributes cross-attention with Gaussian smoothing before reinforcing content-token contributions and attenuating padding contributions, without additional denoiser evaluations. With this intervention, stronger content conditioning can improve prompt alignment at comparable training-image similarity. A local analysis identifies when reinforcement preserves shared value information while redistribution reduces localized attention mass. On Stable Diffusion v1.4 and v2.0, all evaluated smoothing widths lie on the empirical Pareto frontiers for training-image similarity versus both prompt alignment and image preference. A configuration selected on Stable Diffusion reduces template reproduction in DeepFloyd IF without further tuning. These findings support jointly controlling conditioning allocation and strength to generate prompt-consistent alternatives.