CommentsFirst three authors contributed equally. Code are available at https://github.com/amazon-research/mix-generation. Oral presentation at WACV 2023 Pretraining Large Vision and Multimodal Models Workshop
机构
*
University at Buffalo(布法罗大学)
;
NEC Laboratories America(美国 NEC 实验室)
;
Adobe Research(奥多比研究院)
;
Iowa State University(爱荷华州立大学)
;
New York University(纽约大学)
AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering
AV-Master:双路径综合感知实现更优的音频视觉问答
Jiayu Zhang, Shuo Ye, Qilang Ye, Xun Lin, Zihan Song, Zitong Yu
机构
*
Great Bay University(大湾区大学)
;
Dongguan Key Laboratory for Intelligence and Information Technology(东莞智能与信息科技重点实验室)
;
Nankai University(南开大学)
;
Tsinghua Shenzhen International Graduate School(清华大学深圳国际研究生院)
;
The Chinese University of Hong Kong(香港中文大学)
HyLoVQA: Dynamic Hypernetwork-Generated Low-Rank Adaptation for Continual Visual Question Answering
HyLoVQA: 动态超网络生成低秩适应用于连续视觉问答
Yiran Wang, Chenyi Xiong, Ziyue Qin, Miao Zhang, Kui Xiao, Zhifei Li
机构
*
School of Computer Science, Hubei University, Wuhan 430062, China(湖北大学计算机学院,武汉430062,中国)
;
Hubei Key Laboratory of Big Data Intelligent Analysis and Application (Hubei University), Wuhan 430062, China(湖北省大数据智能分析与应用重点实验室(湖北大学),武汉430062,中国)
;
Key Laboratory of Intelligent Sensing System and Security (Hubei University), Ministry of Education, Wuhan 430062, China(智能感知系统与安全重点实验室(湖北大学),教育部,武汉430062,中国)