Text as Any-Modality for Zero-Shot Classification by Consistent Prompt Tuning
机构 * Nanjing University of Science
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV
Comments Accepted for publication at ACMMM 2025
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Nanjing University of Science
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV
Comments Accepted for publication at ACMMM 2025
机构 * Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院) ; Huawei Technologies Co., Ltd.(华为技术有限公司)
专题命中 音频语音多模态 :any-to-any(abstract);分类 cs.AI
Comments Accepted by INTERSPEECH 2025
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments 15 pages, 8 figures, Accepted by ACM MM 2025
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments Published in IEEE Journal of Selected Topics in Signal Processing
Journal ref https://ieeexplore.ieee.org/abstract/document/11086511/
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments 3 pages, 2 tables, submitted for arXiv preprint
专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV
Comments 10 pages, 4 figures
机构 * Google Research Australia(谷歌澳大利亚研究)
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS
Comments arXiv admin note: This paper has been withdrawn by arXiv due to disputed and unverifiable authorship and affiliation
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments 10 Pages, 6 Figures
机构 * Indian Institute of Technology, Madras(印度理工学院马德拉斯分校)
专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV
Comments 6 pages, 1 figure, 1 table. Code and models are released
机构 * Oracle Health & AI(Oracle健康与人工智能)
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI
Comments Accepted to ACL 2025 Industry Track. To appear
机构 * Learning, Adaptive Systems, and Robotics (LASR) Lab, TU Dresden(图腾德斯登技术大学学习、自适应系统与机器人实验室) ; Center for Tactile Internet with Human-in-the-Loop (CeTI)(人机协同触觉互联网中心)
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV
Comments 8 pages. Accepted by 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL
Comments 35 pages, 20 figures
机构 * UNIL ; ETH Zurich(苏黎世联邦理工学院) ; Institute of Neuroinformatics(神经信息学研究所) ; Kivira Health(Kivira健康)
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL
机构 * University of Electronic Science and Technology of China(电子科技大学) ; Chengdu University of Technology(成都理工大学) ; Sun Yat-Sen University(中山大学)
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL
机构 * ETH Zürich(苏黎世联邦理工学院) ; University of Cambridge(剑桥大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
机构 * Cognitive Robotics group, Unit of Automation Technology and Mechanical Engineering, Tampere University(认知机器人组、自动化技术与机械工程单位、塔尔库大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV
Comments Accepted by the 2025 34th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN). Preprint
机构 * Ecole Polytechnique(巴黎高等理工学院) ; MBZUAI(马克斯·普朗克人工智能研究所) ; NTUA(希腊国家技术研究中心) ; University of Peloponnese(希腊皮洛斯大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL
机构 * Tongyi Lab(通义实验室)
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments accepted by Interspeech 2025, 5 pages, 5 tables
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments Accepted to ISMIR 2025
机构 * School of Electronic Science and Engineering(电子科学与工程学院) ; School of Informatics(信息学院)
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments Accepted by InterSpeech 2025
专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS
Comments Submitted to WASAA 2025
机构 * School of Computer Science(计算机科学学院) ; Department of Electronic Engineering(电子工程系) ; School of Cyber Security(网络安全学院)
专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS
Comments Accepted by Interspeech 2025
机构 * School of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(计算机科学与技术学院,南京航空航天大学) ; Urban Data Science Section, Delft University of Technology(都市数据科学部门,代尔夫特理工大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV
机构 * Department of Electrical Engineering and Computer Science, School of Engineering, UQTR, Quebec, Canada(电气工程与计算机科学系,工程学院,UQTR,魁北克,加拿大)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI
Comments Submitted
机构 * Harbin Institute of Technology(哈尔滨工业大学) ; Pengcheng Laboratory(鹏城实验室) ; Shanghai Jiao Tong University(上海交通大学) ; University of Cambridge(剑桥大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL
Comments Accepted in ACL 2025 (Main)
机构 * Emory University(埃默里大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL
机构 * Research Center for Information Technology Innovation(信息技术创新研究中心) ; Institute of Information Science(信息科学研究院)
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments Accepted to Interspeech 2025
专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS
Comments Accepted at ICML 2025 V2: fixed small typo on eq. 15 and eq. 17
机构 * Carnegie Mellon University(卡内基梅隆大学) ; MBZUAI(穆斯林人工智能研究所)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL
Comments ACL 2025 camera-ready