CommentsThis is the author's version of the paper accepted at CHI Conference on Human Factors in Computing Systems (CHI '26), April 13-17, 2026, Barcelona, Spain
CommentsThis is an earlier version of the work released in May 2025. The version accepted at CHI 2026 is available as a separate preprint at arXiv:2511.04366
ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism
Zedong Liu, Shenggan Cheng, Guangming Tan, Yang You, Dingwen Tao
机构
*
Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
;
University of Electronic Science and Technology of China(电子科技大学)
;
National University of Singapore(新加坡国立大学)
UTI-LLM: A Personalized Articulatory-Speech Therapy Assistance System Based on Multimodal Large Language Model
Yudong Yang, Xiaokang Liu, Shaofeng zhao, Rongfeng Su, Nan Yan, Lan Wang
机构
*
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, China(深圳先进技术研究院,中国科学院,中国)
;
Key Laboratory of Biomedical Imaging Science and System, Chinese Academy of Sciences, China(生物医学成像科学与系统重点实验室,中国科学院,中国)
DM-Codec: Distilling Multimodal Representations for Speech Tokenization
Md Mubtasim Ahasan, Md Fahim, Tasnim Mohiuddin, A K M Mahbubur Rahman, Aman Chadha, Tariq Iqbal, M Ashraful Amin, Md Mofijul Islam, Amin Ahsan Ali
机构
*
Center for Computational & Data Sciences, Independent University, Bangladesh(计算与数据科学中心,独立大学,孟加拉国)
;
Amazon GenAI(亚马逊生成人工智能)
;
Qatar Computing Research Institute(卡塔尔计算研究所)
;
University of Virginia(弗吉尼亚大学)
AISHELL6-whisper: A Chinese Mandarin Audio-visual Whisper Speech Dataset with Speech Recognition Baselines
Cancan Li, Fei Su, Juan Liu, Hui Bu, Yulong Wan, Hongbin Suo, Ming Li
机构
*
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
School of Artificial Intelligence, Wuhan University(武汉大学人工智能学院)
;
Beijing AISHELL Technology Co., Ltd.(北京AISHELL科技有限公司)
;
AI Center, OPPO(OPPO人工智能中心)
Visual Authority and the Rhetoric of Health Misinformation: A Multimodal Analysis of Social Media Videos
Mohammad Reza Zarei, Barbara Stead-Coyle, Michael Christensen, Sarah Everts, Majid Komeili
机构
*
School of Computer Science(计算机科学学院)
;
Carleton University(卡尔顿大学)
;
Department of Law and Legal Studies(法律与法律研究系)
;
School of Journalism and Communication(新闻与传播学院)
A Multimodal Symphony: Integrating Taste and Sound through Generative AI
Matteo Spanio, Massimiliano Zampini, Antonio Rodà, Franco Pierucci
机构
*
Centro di Sonologia Computazionale (CSC)(计算声学中心)
;
Department of Information Engineering University of Padova(信息工程系帕多瓦大学)
;
Center for Mind/Brain Sciences (CIMeC)(心智/大脑科学中心)
;
University of Trento(特伦托大学)
;
SoundFood s.r.l.(SoundFood公司)