Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI
Comments Accepted to AAAI 2025
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 音频语音多模态 :audio-visual(title,abstract);multimodal(abstract);分类 cs.CV、cs.AI
Comments Accepted to AAAI 2025
机构 * Faculty of Engineering and Architecture, IDLab-AIRO, Ghent University – imec, Technologiepark 126, 9052 Gent, Belgium(工程与建筑学院,IDLab-AIRO,根特大学–imec,Technologiepark 126,9052 Gent,比利时)
专题命中 音频语音多模态 :multimodal(title,abstract)
专题命中 音频语音多模态 :multimodal(title,abstract)
Comments 18 pages, 5 figures, 4 tables
机构 * Shanghai Jiao Tong University(上海交通大学) ; Shanghai AI Laboratory(上海人工智能实验室) ; Northeastern University(东北大学) ; Carnegie Mellon University(卡内基梅隆大学) ; University of Chinese Academy of Sciences(中国科学院大学) ; Tsinghua University(清华大学) ; Sun Yat-sen University(中山大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS
Comments 26 pages, 23 figures, the code is available at \url{https://github.com/DabDans/AudioMarathon}
机构 * Stanford University(斯坦福大学) ; National University of Singapore(新加坡国立大学) ; Sound Speech and Hearing Clinic(语音与听力诊所)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、cs.AI、eess.AS
Comments EMNLP 2025 Oral Presentation
机构 * McGill University(麦吉尔大学) ; CRBLM ; Mila Quebec AI Institute(魁北克人工智能研究所) ; Université de Montréal(蒙特利尔大学) ; Nouvelle Voix(新声音) ; Montreal Neurological Institute(蒙特利尔神经科学研究所) ; Concordia University(Concordia大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 eess.AS
Comments Accepted to SMASH 2025