Large Language Models for Depression Recognition in Spoken Language Integrating Psychological Knowledge
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL
Journal ref Frontiers in Computer Science, Volume 7, 2025
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL
Journal ref Frontiers in Computer Science, Volume 7, 2025
专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CV
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
机构 * Self-Learning Systems Lab, Faculty of Human Sciences, Department of Special Education and Rehabilitation(自主学习系统实验室,人文科学学院,特殊教育与康复系)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Journal ref Proc. Interspeech 2025
机构 * Project Leader(项目负责人)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV
机构 * Department of Computer Science and Technology, Tsinghua University(计算机科学与技术系,清华大学) ; School of Computer Technology and Application, Qinghai University(计算机技术与应用学院,青海大学) ; ByteDance(字节跳动) ; Nanjing University(南京大学) ; Southern University of Science and Technology(南方科技大学) ; Institute of Acoustics, Chinese Academy of Sciences(中国科学院声学研究所)
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments 34 pages, 10 figures
机构 * Speech and Hearing (SPandH), School of Computer Science, The University of Sheffield, UK(语音与听力(SPandH)、计算机科学学院、谢菲尔德大学、英国) ; South Westphalia University of Applied Sciences, Iserlohn, Germany(西南弗兰肯应用科学大学、伊塞尔洛恩、德国)
专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.MM
Comments Accepted to SPECOM 2025, 13 pages, 4 figures. To appear in the Proceedings of the 27th International Conference on Speech and Computer (SPECOM) 2025, October 13-14, 2025, Szeged, Hungary
机构 * Nanjing University of Science
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV
Comments Accepted for publication at ACMMM 2025
机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) ; Huawei Technologies Co., Ltd.(华为技术有限公司)
专题命中 音频语音多模态 :any-to-any(abstract);分类 cs.AI
Comments Accepted by INTERSPEECH 2025
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments 15 pages, 8 figures, Accepted by ACM MM 2025
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments Published in IEEE Journal of Selected Topics in Signal Processing
Journal ref https://ieeexplore.ieee.org/abstract/document/11086511/
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments 3 pages, 2 tables, submitted for arXiv preprint
专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV
Comments 10 pages, 4 figures
机构 * Google Research Australia(谷歌澳大利亚研究)
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS
Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship and affiliation
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments 10 Pages, 6 Figures
机构 * Indian Institute of Technology, Madras(印度理工学院马德拉斯分校)
专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV
Comments 6 pages, 1 figure, 1 table. Code and models are released
机构 * Oracle Health & AI(Oracle健康与人工智能)
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI
Comments Accepted to ACL 2025 Industry Track. To appear
机构 * Learning, Adaptive Systems, and Robotics (LASR) Lab, TU Dresden(图腾德斯登技术大学学习、自适应系统与机器人实验室) ; Center for Tactile Internet with Human-in-the-Loop (CeTI)(人机协同触觉互联网中心)
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV
Comments 8 pages. Accepted by 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL
Comments 35 pages, 20 figures
机构 * UNIL ; ETH Zurich(苏黎世联邦理工学院) ; Institute of Neuroinformatics(神经信息学研究所) ; Kivira Health(Kivira健康)
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL
机构 * University of Electronic Science and Technology of China(电子科技大学) ; Chengdu University of Technology(成都理工大学) ; Sun Yat-Sen University(中山大学)
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL
机构 * ETH Zürich(苏黎世联邦理工学院) ; University of Cambridge(剑桥大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
机构 * Cognitive Robotics group, Unit of Automation Technology and Mechanical Engineering, Tampere University(认知机器人组、自动化技术与机械工程单位、塔尔库大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV
Comments Accepted by the 2025 34th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN). Preprint
机构 * Ecole Polytechnique(巴黎高等理工学院) ; MBZUAI(马克斯·普朗克人工智能研究所) ; NTUA(希腊国家技术研究中心) ; University of Peloponnese(希腊皮洛斯大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL
机构 * Tongyi Lab(通义实验室)
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments accepted by Interspeech 2025, 5 pages, 5 tables
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments Accepted to ISMIR 2025
机构 * School of Electronic Science and Engineering(电子科学与工程学院) ; School of Informatics(信息学院)
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments Accepted by InterSpeech 2025
专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS
Comments Submitted to WASAA 2025