MuSe 2020 -- The First International Multimodal Sentiment Analysis in Real-life Media Challenge and Workshop
Lukas Stappen, Alice Baird, Georgios Rizos, Panagiotis Tzirakis, Xinchen Du, Felix Hafner, Lea Schumann, Adria Mallol-Ragolta, Björn W. Schuller, Iulia Lefter, Erik Cambria, Ioannis Kompatsiaris
The MuSe 2023 Multimodal Sentiment Analysis Challenge: Mimicked Emotions, Cross-Cultural Humour, and Personalisation
Lukas Christ, Shahin Amiriparian, Alice Baird, Alexander Kathan, Niklas Müller, Steffen Klug, Chris Gagne, Panagiotis Tzirakis, Eva-Maria Meßner, Andreas König, Alan Cowen, Erik Cambria, Björn W. Schuller
Xiaobao Guo, Zitong Yu, Nithish Muthuchamy Selvaraj, Bingquan Shen, Adams Wai-Kin Kong, Alex C. Kot
机构
*
Rapid-Rich Object Search (ROSE) Lab and the College of Computing and Data Science, Nanyang Technological University (NTU)(快速丰富对象搜索(ROSE)实验室和南洋理工大学计算与数据科学学院)
;
School of Computing and Information Technology and Dongguan Key Laboratory for Intelligence and Information Technology, Great Bay University(计算与信息科技学院和东莞智能与信息技术重点实验室,大湾大学)
;
DSO National Laboratories(国防科学实验室)
;
College of Computing and Data Science, Nanyang Technological University (NTU)(计算与数据科学学院,南洋理工大学)
;
SMBU, Shenzhen 518172, China(深圳SMBU,越南河内VinUniversity,和新加坡NTU)
;
VinUniversity, Hanoi 100000, Vietnam
;
and NTU, Singapore
From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs
从感知到决策:多模态大语言模型中听觉与视觉感知的信息流
Wish Suharitdamrong, Muhammad Awais, Xiatian Zhu, Sara Atito
机构
*
Surrey Institute for People-Centred AI (PAI)(萨里人本人工智能研究所)
;
University of Surrey(萨里大学)
;
Centre for Vision, Speech and Signal Processing (CVSSP)(视觉、语音和信号处理中心)
Unified Cross-modal Translation of Score Images, Symbolic Music, and Performance Audio
统一的跨模态评分图像、符号音乐和表演音频翻译
Jongmin Jung, Dongmin Kim, Sihun Lee, Seola Cho, Hyungjoon Soh, Irmak Bukey, Chris Donahue, Dasaem Jeong
机构
*
Department of Artificial Intelligence, Sogang University(西江大学人工智能系)
;
Sogang Future Lab, Sogang University(西江大学未来实验室)
;
Department of Physics Education, Seoul National University(首尔大学物理教育系)
;
Computer Science Department, Carnegie Mellon University(卡内基梅隆大学计算机科学系)
;
Department of Art & Technology, Sogang University(西江大学艺术与技术系)
Taming Modality Entanglement in Continual Audio-Visual Segmentation
驯服持续音频-视觉分割中的模态纠缠
Yuyang Hong, Qi Yang, Tao Zhang, Zili Wang, Zhaojin Fu, Kun Ding, Bin Fan, Shiming Xiang
机构
*
School of Artificial Intelligence, UCAS(人工智能学院,UCAS)
;
MAIS, Institute of Automation(自动化研究所MAIS)
;
School of Intelligent Science and Technology, University of Science and Technolog Beijing(智能科学与技术学院,北京理工大学)
Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
Crab$^{+}$: 一种可扩展且统一的音频视觉场景理解模型,具有显式合作
Dongnuan Cai, Henghui Du, Chang Zhou, Xi Chen, Dan Guo, Hongyuan Zhang, Xuelong Li, Di Hu
机构
*
Gaoling School of Artificial Intelligence, Renmin University of China(中国人民大学耿丽人工智能学院)
;
Institute of Artificial Intelligence of China Telecom (TeleAI)(中国电信人工智能研究院)
;
AI Technology Center, Online Video Business Unit, Tencent PCG(腾讯PCG在线视频业务单元AI技术中心)
;
Hefei University of Technology(合肥工业大学)
;
The University of Hong Kong(香港大学)
A Survey of Generative Categories and Techniques in Multimodal Generative Models
多模态生成模型中生成类别的综述
Longzhen Han, Awes Mubarak, Almas Baimagambetov, Nikolaos Polatidis, Thar Baker
机构
*
School of Architecture, Technology and Engineering, University of Brighton, Lewes Road, BN2 4GJ(建筑、科技与工程学院,布里顿大学,勒斯路,BN2 4GJ)
;
University of Khorfakkan, Sharjah(柯法克坎大学,沙迦)
Audio-Visual Speech Enhancement In Complex Scenarios With Separation And Dereverberation Joint Modeling
Jiarong Du, Zhan Jin, Peijun Yang, Juan Liu, Zhuo Li, Xin Liu, Ming Li
机构
*
School of Cyber Science and Engineering, Wuhan University(武汉大学计算机科学与工程学院)
;
School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)
;
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
Suzhou Municipal Key Laboratory of Multimodal Intelligent Systems, Digital Innovation Research Center, Duke Kunshan University(多模态智能系统苏州市级重点实验室、杜克昆山大学数字创新研究中心)
;
Hardware Engineering System, OPPO(OPPO硬件工程系统)
Audio-Guided Visual Perception for Audio-Visual Navigation
Yi Wang, Yinfeng Yu, Fuchun Sun, Liejun Wang, Wendong Zheng
机构
*
School of Computer Science and Technology, Xinjiang University(新疆大学计算机科学与技术学院)
;
Joint International Research Laboratory of Silk Road Multilingual Cognitive Computing(丝绸之路多语种认知计算国际联合实验室)
;
Tsinghua University(清华大学)
;
Tianjin University of Technology(天津工业大学)
Automating Steering for Safe Multimodal Large Language Models
Lyucheng Wu, Mengru Wang, Ziwen Xu, Tri Cao, Nay Oo, Bryan Hooi, Shumin Deng
机构
*
Zhejiang University(浙江大学)
;
Zhejiang University - Ant Group Joint Lab of Knowledge Graph(浙江大学-蚂蚁集团知识图谱联合实验室)
;
National University of Singapore, NUS-NCS Joint Lab(新加坡国立大学NUS-NCS联合实验室)