Bootstrapping Physics-Grounded Video Generation through VLM-Guided Iterative Self-Refinement
通过VLM引导的迭代自优化提升物理导向的视频生成
Yang Liu, Xilin Zhao, Peisong Wen, Siran Dai, Qingming Huang
机构
*
School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院)
;
School of Computer Science and Technology, Beijing Institute of Technology(北京理工大学计算机科学与技术学院)
;
Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所)
;
School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院)
;
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
机构
*
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
多模态生成式AI:多模态大语言模型、扩散模型与统一
Xin Wang, Yuwei Zhou, Bin Huang, Hong Chen, Wenwu Zhu
机构
*
Department of Computer Science, Beijing Information Science and Technology National Research Center, Tsinghua University(计算机系,北京信息科学与技术国家研究中心,清华大学)