OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset
机构 * 1 Department of Artificial Intelligence, Sogang University, Seoul 04107, Republic of Korea 2 Department of Electronic Engineering, Sogang University, Seoul 04107, Republic of Korea 3 Language Technologies Institute, Carnegie Mellon University, Pittsburgh, PA 15213, USA 4 Mindslab Inc., Gyeonggi-do 13493, Republic of Korea 5 ICT Convergence Disaster/Safety Research Institute, Sogang University, Seoul 04107, Republic of Korea
专题命中 音频语音多模态 :audio-visual(title,abstract);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted to ICASSP 2024