The Sound of Absence: Audio-Language Embedding Models Struggle with Negation
缺失之音:音频-语言嵌入模型在处理否定时存在困难
Chun-Yi Kuan, Hung-yi Lee
机构
*
Graduate Institute of Communication Engineering, National Taiwan University, Taiwan(国家台湾大学通讯工程研究所)
;
Artificial Intelligence Center of Research Excellence (AI-CoRE), National Taiwan University, Taiwan(国家台湾大学人工智能卓越研究中心)
Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach
零努力图像到音乐生成:一种可解释的基于RAG的视觉语言模型方法
Zijian Zhao, Dian Jin, Zijing Zhou
机构
*
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
The Hong Kong Polytechnic University(香港理工大学)
;
The University of Hong Kong(香港大学)
;
The Hong Kong University of Science(香港科学大学)
;
The Hong Kong Polytechnic University Hong Kong China(香港理工大学香港中国)
;
The University of Hong Kong Hong Kong China(香港大学香港中国)
An Approach to Simultaneous Acquisition of Real-Time MRI Video, EEG, and Surface EMG for Articulatory, Brain, and Muscle Activity During Speech Production
一种用于言语产生过程中发音、大脑和肌肉活动的实时MRI视频、脑电图和表面肌电图同步采集方法
Jihwan Lee, Parsa Razmara, Kevin Huang, Sean Foley, Aditya Kommineni, Haley Hsu, Woojae Jeong, Prakash Kumar, Xuan Shi, Yoonjeong Lee, Tiantian Feng, Takfarinas Medani, Ye Tian, Sudarsana Reddy Kadiri, Krishna S. Nayak, Dani Byrd, Louis Goldstein, Richard M. Leahy, Shrikanth Narayanan
机构
*
Signal Analysis and Interpretation Laboratory, University of Southern California(南加州大学信号分析与解释实验室)
;
Ming Hsieh Dept. of Electrical and Computer Engineering, University of Southern California(南加州大学明希斯电气与计算机工程系)
;
Dept. of Linguistics, University of Southern California(南加州大学语言学系)
DeRA-MOS: Optimizing Text-to-Music Evaluation via Decoupled Listwise Ranking and Modality Alignment
DeRA-MOS:通过解耦列表排序和模态对齐优化文本到音乐评估
Chien-Chun Wang, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen
机构
*
E.SUN Financial Holding Co., Ltd.(E.SUN财务控股公司)
;
United Link Co., Ltd.(联合链接有限公司)
;
Institute of Information Science, Academia Sinica(学术院信息科学研究所)
;
Department of Computer Science and Information Engineering, National Taiwan Normal University(台湾师范大学计算机科学与信息工程系)
Towards Conversational Medical AI with Eyes, Ears and a Voice
面向有眼睛、耳朵和声音的对话式医疗AI
Meet Shah, Jason Gusdorf, Anil Palepu, Chunjong Park, Jack W. O'Sullivan, Vishnu Ravi, Tim Strother, Pavel Dubov, Aliya Rysbek, Toshiyuki Fukuzawa, Yana Lunts, Jan Freyberg, Michael B. Chang, Aniruddh Raghu, David Stutz, Devora Berlowitz, Eliseo Papa, Taylan Cemgil, JD Velasquez, Jack Chen, Arthur Chen, Doug Fritz, Charlie Taylor, Katya Tregubova, Jing Rong Lim, Richard Green, Sara Mahdavi, Mahvish Nagda, Jihyeon Lee, Craig Schiff, Liviu Panait, Sukhdeep Singh, Valentin Liévin, David G. T. Barrett, Hannah Gladman, Anna Cupani, Francesca Pietra, Uchechi Okereke, Katherine Tong, Clemens Meyer, Erwan Rolland, Mili Sanwalka, Michael D. Howell, Shixiang Shane Gu, Bibo Xu, Euan A. Ashley, S. M. Ali Eslami, Gregory Wayne, Pushmeet Kohli, Vivek Natarajan, Adam Rodman, Alan Karthikesalingam, Ryutaro Tanno
机构
*
Google DeepMind(谷歌DeepMind)
;
Google Research(谷歌研究)
;
Beth Israel Deaconess Medical Center, Harvard Medical School(贝塞斯达医院, 哈佛医学院)
;
Stanford University(斯坦福大学)
Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models
穿越不确定性:音频感知大语言模型不确定性估计的实证研究
Chun-Yi Kuan, Wei-Ping Huang, Hung-yi Lee
机构
*
Graduate Institute of Communication Engineering, National Taiwan University, Taiwan(台湾大学通讯工程研究所)
;
Artificial Intelligence Center of Research Excellence (AI-CoRE), National Taiwan University, Taiwan(台湾大学人工智能卓越研究中心)
Towards Inclusive Communication: A Unified Framework for Generating Spoken Language from Sign, Lip, and Audio
迈向包容性沟通:一种统一框架,用于从手语、唇形和音频生成口语语言
Jeong Hun Yeo, Hyeongseop Rha, Sungjune Park, Junil Won, Yong Man Ro
机构
*
Integrated Vision Language Laboratory, School of Electrical Engineering, Korea Advanced Institue of Science and Technology (KAIST)(整合视觉语言实验室,电气工程学院,韩国科学技术院(KAIST))