Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
通过VLM引导的迭代自优化提升物理导向的视频生成
Yang Liu, Xilin Zhao, Peisong Wen, Siran Dai, Qingming Huang
机构
*
School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院)
;
School of Computer Science and Technology, Beijing Institute of Technology(北京理工大学计算机科学与技术学院)
;
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
多模态生成式AI:多模态大语言模型、扩散模型与统一
Xin Wang, Yuwei Zhou, Bin Huang, Hong Chen, Wenwu Zhu
机构
*
Department of Computer Science, Beijing Information Science and Technology National Research Center, Tsinghua University(计算机系,北京信息科学与技术国家研究中心,清华大学)
机构
*
Google(谷歌)
;
Tuebingen AI Center/University of Tuebingen(图宾根人工智能中心/图宾根大学)
;
Goethe University Frankfurt(法兰克福歌德大学)
;
MPI for Informatics, Saarland Informatics Campus(信息研究所,萨尔兰信息校区)
;
Technical University of Munich(慕尼黑技术大学)