机构
*
Institute for AI, Peking University(人工智能研究院,北京大学)
;
Beijing Institute for General Artificial Intelligence (BIGAI)(北京通用人工智能研究院)
;
School of Psychological and Cognitive Sciences, Peking University(心理学与认知科学学院,北京大学)
;
School of Computer Science, Peking University(计算机科学学院,北京大学)
;
Yuanpei College, Peking University(元培学院,北京大学)
;
School of Foreign Languages, Peking University(外语学院,北京大学)
;
School of EECS, Peking University(电子工程与科学学院,北京大学)
;
Huazhong University of Science and Technology(华中科技大学)
;
State Key Lab of General AI(通用人工智能国家重点实验室)
;
Nat’l Eng. Research Center of Visual Technology(视觉技术国家工程研究中心)
;
Beijing Key Laboratory of Behavior and Mental Health, Peking University(北京行为与心理健康重点实验室,北京大学)
;
Embodied Intelligence Lab, PKU-Wuhan Institute for Artificial Intelligence(具身智能实验室,北京大学-武汉人工智能研究院)
OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions
OmniVCus: 基于多模态控制条件的前馈主体驱动视频定制
Yuanhao Cai, He Zhang, Xi Chen, Jinbo Xing, Yiwei Hu, Yuqian Zhou, Kai Zhang, Zhifei Zhang, Soo Ye Kim, Tianyu Wang, Yulun Zhang, Xiaokang Yang, Zhe Lin, Alan Yuille
机构
*
Johns Hopkins University(约翰霍普金斯大学)
;
Adobe Research(Adobe研究)
;
The University of Hong Kong(香港大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Shanghai Jiao Tong University(上海交通大学)
专题命中
视频多模态
:multimodal(title);分类 cs.CV
AI总结
OmniVCus通过多模态控制条件和改进的嵌入机制实现高效的多主体视频定制。
CommentsNeurIPS 2025; A data construction pipeline and a diffusion Transformer framework for controllable subject-driven video customization
机构
*
Shanghai AI Laboratory(上海人工智能实验室)
;
Nanjing University(南京大学)
;
Shanghai Jiao Tong University(上海交通大学)
;
The Chinese University of Hong Kong(香港中文大学)
机构
*
School of Electronic Information and Communications, Huazhong University of Science and Technology(电子信息与通信学院,华中科技大学)
;
Peng Cheng Laboratory(鹏城实验室)
;
Pazhou Laboratory (Huangpu)(琶洲实验室(黄埔))
;
School of Mechanical Engineering and Electronic Information, China University of Geosciences(机械工程与电子信息学院,中国地质大学)
;
State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications(网络与交换技术国家重点实验室,北京邮电大学)
MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
MM-SpuBench:迈向更好地理解多模态大语言模型中伪偏差的深入研究
Wenqian Ye, Bohan Liu, Guangtao Zheng, Di Wang, Yunsheng Ma, Xu Cao, Bolin Lai, James M. Rehg, Aidong Zhang
机构
*
University of Virginia(弗吉尼亚大学)
;
Purdue University(普渡大学)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Georgia Institute of Technology(佐治亚理工学院)
机构
*
University of Manchester(曼彻斯特大学)
;
Queen Mary University of London(伦敦大学玛丽女王学院)
;
Hongkong University of Science and Technology(香港科学与技术大学)
;
Nanjing University(南京大学)
;
Dartmouth College(达特茅斯学院)
机构
*
Shanghai Jiao Tong University(上海交通大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Fudan University(复旦大学)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))