SALAD-VAE: Semantic Audio Compression with Language-Audio Distillation
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments submitted to ICASSP 2026
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments submitted to ICASSP 2026
机构 * Bar Ilan University(巴伊兰大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI
机构 * McGill University(麦吉尔大学) ; CRBLM ; Mila Quebec AI Institute(魁北克人工智能研究所) ; Université de Montréal(蒙特利尔大学) ; Nouvelle Voix(新声音) ; Montreal Neurological Institute(蒙特利尔神经科学研究所) ; Concordia University(Concordia大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments Accepted to SMASH 2025
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL
Comments 11 pages, 4 figures, 3 tables
机构 * Singapore University of Technology and Design(新加坡科技设计大学) ; Hochschule für Musik Detmold(音乐学院Detmold)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI
机构 * Georgia Institute of Technology(佐治亚理工学院) ; Carnegie Mellon University(卡内基梅隆大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV
Comments ICCV 2025. Project page: https://clink-chop-thud.github.io/
专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS
机构 * Georgia Institute of Technology(佐治亚理工学院) ; Bytedance Inc.(字节跳动公司) ; Squirrel AI, USA(squirrel AI 美国分公司) ; The University of Virginia(弗吉尼亚大学) ; Cornell University(康奈尔大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV
Comments Github Repo: https://github.com/AdityaLab/MM4TSA Updated to include papers accepted by IJCAI25, KDD25, ICML25, NeurIPS25 4 figures or tables, 19 pages, 251 references
机构 * School of Computer Science, University of Sheffield(谢菲尔德大学计算机科学学院) ; Centre for Infectious Disease Research in Zambia(赞比亚传染病疾病研究中心)
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments submitted to IEEE Journal of Biomedical and Health Informatics
机构 * Tampere University, Tampere, Finland(塔尔库大学) ; University of Oxford, Oxford, UK(牛津大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV
Comments Preprint version. The Version of Record is published in DAGM GCPR 2025 proceedings with Springer Lecture Notes in Computer Science (LNCS). Updated results and resources are available at the project page: https://saganet.notion.site
机构 * School of Artificial Intelligence and Computer Science(人工智能与计算机科学学院) ; Institute of Acoustics Chinese Academy of Science(中国科学院声学研究所)
专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV
Comments Submitted to ICASSP 2026
机构 * Kyutai
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL
机构 * Meta Reality Labs Research(Meta现实实验室)
专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV
Comments Accepted to IJCV
机构 * Tsinghua University(清华大学) ; Peking University(北京大学) ; AMAP Speech(AMAP语音)
专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS
Comments conference paper about TTS
机构 * ByteDance(字节跳动)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV
Comments Project Page at https://byteaigc.github.io/X-Streamer
机构 * University of Chinese Academy of Sciences(中国科学院大学) ; State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院人工智能安全国家重点实验室,计算技术研究所)
专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV
机构 * University of Science and Technology of China(中国科学技术大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; Sichuan University(四川大学) ; Shanghai Jiao Tong University(上海交通大学) ; University of Macau(澳门大学) ; Tsinghua University(清华大学) ; The Chinese University of Hong Kong(香港中文大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL
Comments 34 pages
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI
Comments The version accepted for publication at International Journal of Human-Computer Studies
Journal ref International Journal of Human-Computer Studies (2025)
机构 * INSPER Institute of Teaching and Research(INSPER教学与研究机构) ; Research University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校研究大学)
专题命中 音频语音多模态 :cross-modal(abstract);分类 eess.AS
机构 * Hong Kong University of Science and Technology(香港理工大学) ; Soul AI Lab(Soul AI 实验室)
专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.AI
专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS
Comments Proceedings of Interspeech 2025
机构 * Guangdong Provincial Key Laboratory of Ultra High Definition Immersive Media Technology, Peking University, Shenzhen(广东超高清沉浸媒体技术重点实验室,北京大学深圳校区) ; X-LANCE Lab, Shanghai Jiao Tong University, Shanghai(X-LANCE实验室,上海交通大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
机构 * ADAPT Centre(ADAPT中心) ; Dublin City University (DCU)(都柏林城市大学) ; Trinity College Dublin (TCD)(三一学院都柏林)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL
Comments Accepted at WMT2025 (ENNLP) for oral presented
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI
机构 * IEEE
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV
Comments 5 pages, 4 figures,
机构 * University of Science and Technology Beijing(北京科技大学) ; Tencent AI Lab(腾讯AI实验室) ; National University of Singapore(新加坡国立大学)
专题命中 音频语音多模态 :multi-modal(abstract);分类 eess.AS
Comments Accepted by APSIPA ASC2025
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI
Comments Companion website for additional chapters: https://www.sscardapane.it/alice-book
机构 * Singapore Institute of Technology(新加坡理工学院) ; Institute Of Acoustics, Chinese Academy Of Sciences(中国科学院声学研究所) ; Duke Kunshan University(杜克大学昆山分校)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI
Comments The paper has been accepted by APCIPA ASC 2025
专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS
Comments 5 pages
机构 * Department of Electrical Engineering(电气工程系) ; Department of Semiconductor Engineering(半导体工程系) ; Center for Semiconductor Technology Convergence(半导体技术融合中心)
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.AI
Journal ref Interspeech 2025