机构
*
Australian Institute for Machine Learning, University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学)
;
Sony Group Corporation(索尼集团)
;
Sony AI, Sony Group Corporation(索尼人工智能,索尼集团)
Predicting Brain Responses To Natural Movies With Multimodal LLMs
Cesar Kadir Torrico Villanueva, Jiaxin Cindy Tu, Mihir Tripathy, Connor Lane, Rishab Iyer, Paul S. Scotti
机构
*
Medical AI Research Center (MedARC)(医学人工智能研究中心(MedARC))
;
Psychological and Brain Sciences, Dartmouth College(心理学与脑科学系,达特茅斯学院)
;
Core for Advanced Magnetic Resonance Imaging (CAMRI), Baylor College of Medicine(先进磁共振成像核心(CAMRI),贝勒医学院)
;
Sophont
;
Princeton Neuroscience Institute(普林斯顿神经科学研究所)
Versatile Multimodal Controls for Expressive Talking Human Animation
Zheng Qin, Ruobing Zheng, Yabing Wang, Tianqi Li, Zixin Zhu, Sanping Zhou, Ming Yang, Le Wang
机构
*
National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Xi'an Jiaotong University, Ant Group(人机混合增强智能国家级实验室,西安交通大学,蚂蚁集团)
;
National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, Xi'an Jiaotong University(人机混合增强智能国家级实验室,西安交通大学)
;
University at Buffalo(布法罗大学)
Speech as a Multimodal Digital Phenotype for Multi-Task LLM-based Mental Health Prediction
Mai Ali, Christopher Lucasius, Tanmay P. Patel, Madison Aitken, Jacob Vorstman, Peter Szatmari, Marco Battaglia, Deepa Kundur
机构
*
The Edward S. Rogers Sr. Department of Electrical and Computer Engineering, University of Toronto, Toronto, Canada(电气与计算机工程系,多伦多大学)
;
Division of Engineering Science, University of Toronto, Toronto, Canada(工程科学系,多伦多大学)
;
Cundill Centre for Child and Youth Depression, Centre for Addiction and Mental Health, Toronto, Canada(儿童与青少年抑郁研究中心,成瘾与心理健康中心)
;
Department of Psychology, York University, Toronto, Canada(心理学系,约克大学)
;
The Hospital for Sick Children, Toronto, ON, Canada(多伦多儿童医院)
;
Department of Psychiatry, University of Toronto, Toronto, Canada(精神病学系,多伦多大学)
Comments6 pages, 1 figure, 3 tables. The corresponding author is Mai Ali (maia dot ali at mail dot utoronto dot ca). Christopher Lucasius and Tanmay P. Patel contributed equally
Exploring Human-AI Complementarity in CPS Diagnosis Using Unimodal and Multimodal BERT Models
Kester Wong, Sahan Bulathwela, Mutlu Cukurova
机构
*
UCL Knowledge Lab, Institute of Education, University College London, UK(伦敦大学学院教育研究所知识实验室)
;
UCL Centre for Artificial Intelligence, Department of Computer Science, University College London, UK(伦敦大学学院人工智能中心计算机科学系)
CommentsAccepted to appear in the workshop proceedings for the HEXED'25 workshop in the 26th International Conference on Artificial Intelligence in Education 2025 (AIED 2025), 22 July 2025, Palermo, Italy. 5 pages
CorMulT: A Semi-supervised Modality Correlation-aware Multimodal Transformer for Sentiment Analysis
Yangmin Li, Ruiqi Zhu, Wengen Li
机构
*
Information Networking Institute at Carnegie Mellon University (CMU)(卡内基梅隆大学信息网络研究所)
;
School of Computer Science and Technology at Tongji University(同济大学计算机科学与技术学院)
SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation
Youliang Zhang, Zhaoyang Li, Duomin Wang, Jiahe Zhang, Deyu Zhou, Zixin Yin, Xili Dai, Gang Yu, Xiu Li
机构
*
Tsinghua University(清华大学)
;
StepFun
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
;
The Hong Kong University of Science and Technology(香港科技大学)
Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization
Md Moinul Islam, Sofoklis Kakouros, Janne Heikkilä, Mourad Oussalah
机构
*
Center for Machine Vision and Signal Analysis, Faculty of ITEE, University of Oulu, Finland(机器视觉与信号分析中心,ITEE学院,奥卢大学,芬兰)
;
Center for Machine Vision(机器视觉与信号分析中心)
;
Signal Analysis, Faculty of ITEE, University of Oulu, Finland(信号分析,ITEE学院,奥卢大学,芬兰)
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction
Wenxuan Wu, Shuai Wang, Xixin Wu, Helen Meng, Haizhou Li
机构
*
Department of SEEM(SEEM系)
;
SRIBD, School of Data Science(数据科学学院)
;
School of Intelligence Science and Technology(智能科学与技术学院)
;
Department of ECE(电子工程系)
MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge
Xin Jing, Jiadong Wang, Iosif Tsangko, Andreas Triantafyllopoulos, Björn W. Schuller
机构
*
Technical University of Munich(慕尼黑技术大学)
;
Imperial College London(伦敦帝国理工学院)
;
CHI – Chair of Health Informatics(健康信息学系)
;
Munich Centre for Machine Learning(慕尼黑机器学习中心)
;
GLAM – Group on Language, Audio, & Music(语言、音频与音乐小组)
Steganography Beyond Space-Time with Chain of Multimodal AI
Ching-Chun Chang, Isao Echizen
机构
*
Information and Society Research Division, National Institute of Informatics(信息与社会研究部,日本信息机构)
;
Graduate School of Information Science and Technology, University of Tokyo(东京大学信息科学与技术研究生院)
;
School of Multidisciplinary Sciences, Graduate University for Advanced Studies (SOKENDAI)(多学科科学学院,研究生高级研究大学(SOKENDAI))