arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ToPO:面向注意力型潜在扩散模型的令牌条件偏好路由

ToPO: Token-Conditioned Preference Routing for Attention-Based Latent Diffusion Models

Juntao Xu, Shihong Li, Hoi Fan Au, Ning Zhu

arXiv 2609.03688首次发表:更新:

发表机构

Tsinghua University; University of Electronic Science and Technology of China; Stanford University(清华大学; 电子科技大学; 斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出ToPO方法,通过构建时空路由优化注意力型潜在扩散模型的偏好,在SD-1.5和SDXL相关指标上优于Diffusion-DPO,且在盲测中胜率更高。

AI 中文摘要

成对偏好标签可对完整图像进行排序,但Diffusion-DPO会将其作用施加于大量空间和去噪时间坐标。针对基于注意力、噪声预测的潜在扩散模型,ToPO(面向令牌的偏好优化)从冻结参考去噪器中的分支平方残差对比,构建每个小批次的、分离的、可拆分的时空路由。偏好分支的交叉注意力使用内容令牌调制空间因子,且在无局部标签或学习奖励模型的情况下,添加辅助像素中点排序项。在采用相同更新计划的三次匹配种子重训练中,ToPO在所有5项已报告的SD-1.5指标上,以及SDXL的HPSv2、ImageReward和CLIP指标上,均比Diffusion-DPO具有更高的端点估计值;在汇总的盲测SDXL A/B研究中,其原始胜率也更高。这些发现仅适用于所报告的等更新U-Net协议,而非等计算量对比。

英文摘要

Pairwise preference labels rank complete images, yet Diffusion-DPO applies their effect over many spatial and denoising-time coordinates. For attention-based, noise-prediction latent diffusion, ToPO (Token-Oriented Preference Optimization) constructs a per-minibatch, detached, separable spatial-temporal route from branchwise squared-residual contrast in a frozen reference denoiser. Preferred-branch cross-attention uses content tokens to modulate the spatial factor, and an auxiliary pixel-midpoint ordering term is added without local labels or a learned reward model. In matched three-seed retrainings with a shared update schedule, ToPO has higher endpoint estimates than Diffusion-DPO on all five reported SD-1.5 metrics and on HPSv2, ImageReward, and CLIP for SDXL. It also receives larger raw win shares in an aggregate blind SDXL A/B study. These findings are scoped to the reported equal-update U-Net protocols rather than an equal-compute comparison.

Comments37 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑