arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

期刊&会议

IEEE TPAMI

IEEE Transactions on Pattern Analysis and Machine Intelligence · 期刊 · Computer Vision

2026-03-02 至 2026-03-02 共收录 2
2411.11727 2026-03-02 cs.LG cs.CV

Aligning Few-Step Diffusion Models with Dense Reward Difference Learning

对齐少步扩散模型与密集奖励差学习

Ziyi Zhang, Li Shen, Sen Zhang, Deheng Ye, Yong Luo, Miaojing Shi, Dongjing Shan, Bo Du, Dacheng Tao

机构 * School of Computer Science, National Engineering Research Center for Multimedia Software and Hubei Key Laboratory of Multimedia and Network Communication Engineering, Wuhan University(计算机学院、多媒体软件国家工程研究中心和多媒体与网络通信工程湖北省重点实验室、武汉大学) School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University(网络安全科学与技术学院、中山大学深圳校区) TikTok, ByteDance(TikTok、字节跳动) Tencent Inc.(腾讯公司) College of Electronic and Information Engineering, Tongji University(电子信息工程学院、同济大学) School of Medical Information and Engineering, Southwest Medical University(医学信息与工程学院、西南医科大学) College of Computing and Data Science and the Generative AI Lab at Nanyang Technological University(计算与数据科学学院和南洋理工大学生成式AI实验室)

AI总结 SDPO通过双状态轨迹采样和密集奖励差学习,提升少步扩散模型在低步数下的对齐性能和优化效率。

Comments Accepted by IEEE TPAMI

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.23836 2026-03-02 cs.CV

Stereo-Talker: Audio-driven 3D Human Synthesis with Prior-Guided Mixture-of-Experts

立体说话者:基于优先级的混合专家引导的音频驱动3D人类合成

Xiang Deng, Youxin Pang, Xiaochen Zhao, Chao Xu, Lizhen Wang, Hongjiang Xiao, Shi Yan, Hongwen Zhang, Yebin Liu

机构 * Department of Automation, Tsinghua University(自动化系,清华大学) Bytedance Inc.(字节跳动公司) State Key Laboratory of Media Convergence and Communication, Communication University of China(媒体融合与传播国家重点实验室,中国传媒大学) School of Artificial Intelligence, Beijing Normal University(人工智能学院,北京师范大学)

AI总结 Stereo-Talker通过优先级引导的混合专家机制实现音频驱动的3D人类合成,生成具有精确唇同步和逼真质量的视频。

Journal ref TPAMI 2025

详情

展开后加载摘要…

URL PDF HTML 收藏