发表机构
Mohamed bin Zayed University of Artificial Intelligence; United Arab Emirates University(穆罕默德·本·扎耶德人工智能大学; 阿联酋大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对扩散编辑模型参考注意力分配不足的问题,提出无需训练的RefGAP方法,在线校正logit偏移以增强参考使用,在七个编辑模型上提升身份保真度并保持良好权衡。
AI 中文摘要
参考引导的扩散编辑模型难以忠实再现用户提供的参考。我们识别出扩散编辑模型中的一个潜在瓶颈:许多方法提供的参考注意力分配有限。例如,在LoomVideo中,编辑区域查询分配给参考的注意力质量不足1%。我们提出RefGAP,一种无需训练的校正方法,在前向传播过程中根据测量的参考注意力质量,在线确定每层的logit偏移幅度。对参考logit的正偏移增强编辑区域查询对参考的使用,而对保持区域查询的负偏移限制编辑区域外由参考引起的改变。两个全局系数控制校正;它们从四个开发扩散编辑模型的验证数据中一次性选择并保持固定。在七个基于扩散的图像/视频编辑模型中,RefGAP提高了头部交换和面部交换中的身份保真度。RefGAP实现了与单独调优的恒定编辑侧偏置相当的保真度-保留权衡,无需逐方法强度扫描。在虚拟试穿和背景替换上的额外实验评估了超越身份编辑的迁移能力。
英文摘要
Reference-guided diffusion editors struggle to faithfully reproduce user-provided references. We identify a potential bottleneck in diffusion editors: many methods provide limited reference-attention allocation. For example, in LoomVideo, edit-region queries assign less than 1% of their attention mass to the reference. We introduce RefGAP, a training-free correction that determines logit-offset magnitudes online at each layer from the reference-attention mass measured during the forward pass. Positive offsets to reference logits strengthen reference usage by edit-region queries, while negative offsets for keep-region queries limit reference-induced changes outside the edit. Two global coefficients control the correction; they are selected once on validation data from four development diffusion editors and held fixed. Across seven diffusion-based image/video editors, RefGAP improves identity fidelity in head swapping and face swapping. RefGAP achieves a fidelity-preservation trade-off comparable to separately tuned constant edit-side biases, without per-approach strength sweeps. Additional experiments on virtual try-on and background replacement evaluate transfer beyond identity editing.