Aligning Few-Step Diffusion Models with Dense Reward Difference Learning
对齐少步扩散模型与密集奖励差学习
机构 * School of Computer Science, National Engineering Research Center for Multimedia Software and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University(计算机学院、多媒体软件国家工程研究中心和多媒体与网络通信工程湖北省重点实验室、武汉大学) ; School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(网络安全科学与技术学院、中山大学深圳校区) ; TikTok, ByteDance(TikTok、字节跳动) ; Tencent Inc.(腾讯公司) ; College of Electronic and Information Engineering, Tongji University(电子信息工程学院、同济大学) ; School of Medical Information and Engineering, Southwest Medical University(医学信息与工程学院、西南医科大学) ; College of Computing and Data Science and the Generative AI Lab at Nanyang Technological University(计算与数据科学学院和南洋理工大学生成式AI实验室)
AI总结 SDPO通过双状态轨迹采样和密集奖励差学习,提升少步扩散模型在低步数下的对齐性能和优化效率。
Comments Accepted by IEEE TPAMI
Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026