Audio-Enhanced Vision-Language Modeling with Latent Space Broadening for High Quality Data Expansion
机构 * ByteDance Inc.(字节跳动公司)
专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * ByteDance Inc.(字节跳动公司)
专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM
机构 * CASIA(中国科学院自动化研究所) ; UCAS(中国科学技术大学) ; ZGCA(北京智感科技有限公司) ; HKU(香港大学) ; HKUST(香港科技大学) ; NTU(国立台湾大学) ; PKU(北京大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
Comments Preprint, Under review
专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)
Comments 10 pages
专题命中 音频语音多模态 :multimodal(title,abstract);cross-modal(abstract)
机构 * Center for Robotics, Mines Paris - PSL University Paris, France(机器人中心,巴黎 Mines Paris - PSL 大学)
专题命中 音频语音多模态 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV、eess.AS
Comments ICLR 2025 (Poster). Camera ready version. Project Page: https://amandinebtto.github.io/NeRAF; 24 pages, 13 figures
Journal ref The Thirteenth International Conference on Learning Representations, 2025
机构 * Georgia Institute of Technology(佐治亚理工学院) ; Carnegie Mellon University(卡内基梅隆大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV
Comments ICCV 2025. Project page: https://clink-chop-thud.github.io/
专题命中 音频语音多模态 :audio-visual(abstract);分类 eess.AS
机构 * Georgia Institute of Technology(佐治亚理工学院) ; Bytedance Inc.(字节跳动公司) ; Squirrel AI, USA(squirrel AI 美国分公司) ; The University of Virginia(弗吉尼亚大学) ; Cornell University(康奈尔大学)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV
Comments Github Repo: https://github.com/AdityaLab/MM4TSA Updated to include papers accepted by IJCAI25, KDD25, ICML25, NeurIPS25 4 figures or tables, 19 pages, 251 references
机构 * MSc Artificial Intelligence Master Thesis(人工智能硕士论文)
专题命中 音频语音多模态 :multi-modal(abstract)
机构 * University Of Southern California(美国南加州大学)
专题命中 视频多模态 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.CL、cs.AI
机构 * MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
机构 * Department of Software \& Microelectronics Peking University Beijing, China ; Department of Software \& Microelectronics Peking University Beijing, China hangli\ ; Department of Software \& Microelectronics Peking University Beijing, China zehua\ ; Department of Software \& Microelectronics Peking University Beijing, China xiaofan\ ; Economic Law School China University of Political Science ; Peking University Changsha, China
专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CL、cs.AI
机构 * Department of Electrical and Electronic Engineering Imperial College London(帝国理工学院电子与电气工程系) ; Department of Civil and Environmental Engineering Imperial College London(帝国理工学院土木与环境工程系) ; Institute of Biomedical Engineering Department of Engineering Science University of Oxford(牛津大学生物医学工程研究所)
专题命中 跨模态检索 :multimodal(title)
Comments 4 pages, 2 figures. Accepted for oral presentation at the 52nd international Computing in Cardiology Conference (CinC2025)
机构 * Character AI ; Yale University(耶鲁大学)
专题命中 多模态生成 :cross-modal(title,abstract);multimodal(abstract);分类 cs.CV、cs.MM、eess.AS
专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL
Comments This paper is accepted for presentation in TRB annual meeting 2026. The version presented here is the preprint version before peer review process
机构 * National University of Singapore(新加坡国立大学) ; Shanghai Jiao Tong University(上海交通大学) ; Shanghai AI Lab(上海人工智能实验室)
专题命中 多模态生成 :multi-modal(title)
机构 * School of Vehicle and Mobility & College of AI, Tsinghua University(车辆与移动学院及人工智能学院,清华大学) ; School of Artificial Intelligence, University of Chinese Academy of Sciences(人工智能学院,中国科学院大学) ; School of Mechanical Engineering, University of Science and Technology Beijing(机械工程学院,北京科技大学)
专题命中 多模态生成 :multimodal(abstract);分类 cs.AI
专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI
Comments Fixed and extended results
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV
Comments NeurIPS 2024. Project page at https://diffcut-segmentation.github.io. Code at https://github.com/PaulCouairon/DiffCut
专题命中 多模态生成 :multimodal(abstract)
Comments 14 pages, to be published at the 26th International Conference on Artificial Intelligence in Education (AIED '25)
专题命中 多模态评测 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI
Comments 7 pages, 2 figure, 2 tables, CV4A11y Workshop at ICCV 2025
机构 * AWS AI Labs(AWS人工智能实验室) ; Language Technology Lab, University of Cambridge(语言技术实验室,剑桥大学)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI
Comments Accepted by EMNLP 2025
机构 * Institute of Biomedical Image Analysis UMIT TIROL -- Private University for Health Sciences(生物医学影像分析研究所)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments Contribution to Conference for Computer Assisted Radiology and Surgery (CARS 2025)
Journal ref Int J CARS (2025)
机构 * Argonne National Laboratory(阿贡国家实验室)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
Comments Preprint
专题命中 多模态评测 :multi-modal(title,abstract)
Comments 39 pages, 13 figures, Supporting Information available in SI.pdf
机构 * School of Computer Science & Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) ; School of Computer Science, University of Nottingham Ningbo China(诺丁汉大学宁波校区计算机科学学院) ; Wenzhou Medical University(温州医科大学) ; School of Computer Science, University of Nottingham(诺丁汉大学计算机科学学院)
专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CV
Comments Strong accept by NeurIPS2025 Reviewers and AC
机构 * University of Oxford(牛津大学)
专题命中 多模态评测 :multimodal(abstract);分类 cs.CV、cs.AI
机构 * Indian Institute of Technology Patna(印度理工学院帕纳布分校) ; Sardar Patel Institute of Technology(萨达尔·帕特尔技术学院) ; Universitas Gadjah Mada(加查马大学) ; King Mongkut’s Institute of Technology Ladkrabang(拉差班国王技术学院) ; Shenzhen Technology University(深圳技术大学) ; Université de Toulouse(图卢兹大学)
专题命中 多模态评测 :multimodal(abstract);分类 cs.CL、cs.AI
Comments 52 pages, 56 figures; appearing at EMNLP'25
机构 * Fudan University(复旦大学) ; Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; Imperial College London(帝国理工学院) ; University of Cambridge(剑桥大学)
专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV
Comments 26 pages, 13 figures
机构 * ECE Department, University of Alabama in Huntsville(阿拉巴马大学亨茨维尔分校电子与计算机工程系)
专题命中 多模态Agent :multimodal(title,abstract);分类 cs.AI
Comments Accepted at IEEE Transactions on Network Science and Engineering