机构
*
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
Prometheus Vision Technology Co., Ltd.(普罗米修斯视觉科技有限公司)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
专题命中
音频语音多模态
:multi-modal(abstract);分类 cs.CV
CommentsProceedings of the Computer Graphics International 2025 (CGI'25)
CommentsThis article will serve as an extension of the preceding work, "VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models" (arXiv:2505.15727). Therefore, we have chosen to withdraw to avoid potential duplicate publication. We will update the previously open-sourced paper of VocalBench in several weeks to include the content of VocalBench-zh
MERaLiON-SER: Robust Speech Emotion Recognition Model for English and SEA Languages
Hardik B. Sailor, Aw Ai Ti, Chen Fang Yih Nancy, Chiu Ying Lay, Ding Yang, He Yingxu, Jiang Ridong, Li Jingtao, Liao Jingyi, Liu Zhuohan, Lu Yanfeng, Ma Yi, Manas Gupta, Muhammad Huzaifah Bin Md Shahrin, Nabilah Binte Md Johan, Nattadaporn Lertcheva, Pan Chunlei, Pham Minh Duc, Siti Maryam Binte Ahmad Subaidi, Siti Umairah Binte Mohammad Salleh, Sun Shuo, Tarun Kumar Vangani, Wang Qiongqiong, Won Cheng Yi Lewis, Wong Heng Meng Jeremy, Wu Jinyang, Zhang Huayun, Zhang Longyin, Zou Xunlong
机构
*
MERaLiON Team Institute for Infocomm Research (I 2 R), A*STAR, Singapore(MERaLiON团队信息与通信研究所(I 2 R),A*STAR,新加坡)
Visual Cues Enhance Predictive Turn-Taking for Two-Party Human Interaction
Sam O'Connor Russell, Naomi Harte
机构
*
ADAPT Centre, School of Engineering, Trinity College Dublin(ADAPT中心、工程学院、都柏林信任学院)
专题命中
音频语音多模态
:multimodal(abstract);分类 cs.CL
CommentsAccepted to ACL 2025, Findings of the Association for Computational Linguistics
Journal refIn Findings of the Association for Computational Linguistics: ACL 2025, pages 209--221, Vienna, Austria. Association for Computational Linguistics, 10.18653/v1/2025.findings-acl.12
Deep Insights into Cognitive Decline: A Survey of Leveraging Non-Intrusive Modalities with Deep Learning Techniques
David Ortiz-Perez, Manuel Benavent-Lledo, Jose Garcia-Rodriguez, David Tomás, M. Flores Vizcaya-Moreno
机构
*
Dept. of Computer Science and Technology(计算机科学与技术系)
;
University of Alicante(阿尔瓦登特大学)
;
Unit of Clinical Nursing Research(临床护理研究单位)
;
Faculty of Health Sciences(健康科学学院)
Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges
Cheng Huang, Nyima Tashi, Fan Gao, Yutong Liu, Jiahao Li, Hao Tian, Siyang Jiang, Thupten Tsering, Ban Ma-bao, Renzeg Duojie, Gadeng Luosang, Rinchen Dongrub, Dorje Tashi, Jin Zhang, Xiao Feng, Hao Wang, Jie Tang, Guojie Tang, Xiangxiang Wang, Jia Zhang, Tsengdar Lee, Yongbin Yu
机构
*
University of Electronic Science and Technology of China(电子科技大学)
;
Southern Methodist University(南方 Methodist 大学)
;
The City University of Hong Kong(香港城市大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
University of Connecticut(康涅狄格大学)
;
Tsinghua University(清华大学)
;
University of Texas at Arlington(德克萨斯大学阿灵顿分校)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Tibet University(西藏大学)
PASE: Phoneme-Aware Speech Encoder to Improve Lip Sync Accuracy for Talking Head Synthesis
Yihuan Huang, Jiajun Liu, Yanzhen Ren, Jun Xue, Wuyang Liu, Zongkun Sun
机构
*
Key Laboratory of Aerospace Information Security and Trusted Computing, Ministry of Education, School of Cyber Science and Engineering, Wuhan University(航空航天信息安全部分和可信计算重点实验室、教育部、网络安全科学与工程学院、武汉大学)
;
School of Cyber Science and Engineering, Wuhan University(网络安全科学与工程学院、武汉大学)
;
School of Police Information, Shandong Police College(警务信息学院、山东警察学院)
Chord Colourizer: A Near Real-Time System for Visualizing Musical Key
Paul Haimes
机构
*
Ritsumeikan University(立命馆大学)
专题命中
音频语音多模态
:multi-modal(abstract);分类 eess.AS
CommentsAuthor copy. This paper is in press for presentation at ADADA 2025. Please cite as: Haimes, P. (in press). Chord Colourizer: A near real-time system for visualizing musical key. In Proceedings of the 23rd International Conference of Asia Digital Art and Design Association (ADADA)