Low Resource Multimodal Translation of Nepali Spoken Words into Emotion-Conditioned Sign Language Avatars
低资源场景下尼泊尔口语词汇到情感条件手语虚拟人物的多模态翻译
Jatin Bhusal, Salma Tamang
机构
*
Center for Human Mobility and Communications, Prateek Innovations(普拉蒂克创新公司人类移动与通信中心)
;
Sunway International Business School, Birmingham City University(双威国际商学院,伯明翰城市大学)
机构
*
Kahlert School of Computing, University of Utah(犹他大学卡勒特计算学院)
;
Scientific Computing and Imaging Institute, University of Utah(犹他大学科学计算与成像研究所)
;
Department of Electrical and Computer Engineering, University of Utah(犹他大学电气与计算机工程系)
XAI-Grounded Explanation Generation for Speech Deepfake Detection with Training-Free Multimodal Large Language Models
基于XAI的语音深度伪造检测解释生成:使用免训练多模态大语言模型
Yupei Li, Qiyang Sun, Xiaoliang Wu, Chenxi Wang, Berrak Sisman, Björn W. Schuller
机构
*
Imperial College London(帝国理工学院)
;
Technical University of Munich(慕尼黑工业大学)
;
University of Southampton(南安普顿大学)
;
MBZUAI(穆罕默德·本·扎耶德人工智能大学)
;
Johns Hopkins University(约翰霍普金斯大学)
Cross-Modal Masking for Robust Silent Speech Synthesis Using sEMG and Lipreading
基于sEMG和唇读的鲁棒无声语音合成的跨模态掩蔽
Eder del Blanco, David Gimeno-Gómez, Eva Navas, Carlos-D. Martínez-Hinarejos, Inma Hernáez
机构
*
Aholab research group within the HiTZ Center at University of the Basque Country (UPV/EHU)(巴斯克大学HiTZ中心内Aholab研究组)
;
PRHLT research center, Universitat Politècnica de València (UPV)(瓦伦西亚理工大学PRHLT研究中心)