Multimodal ML: Quantifying the Improvement of Calorie Estimation Through Image-Text Pairs
专题命中 图文多模态 :multimodal(title,abstract);image-text(title);分类 cs.CV
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 图文多模态 :multimodal(title,abstract);image-text(title);分类 cs.CV
专题命中 图文多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI
Comments Accepted at AAAI 2026
机构 * Peng Zhang 1 , Zhihui Lai 1 , Wenting Chen 2 1 1 footnotemark: 1 , Xu Wu 1 , Heng Kong 3(某机构)
专题命中 图文多模态 :cross-modal(title,abstract);image-text(abstract);分类 cs.CV
Comments AAAI 2026
机构 * New York University(纽约大学) ; Cornell Tech(康奈尔科技) ; Vanderbilt University(范德比大学) ; Carnegie Mellon University(卡内基梅隆大学) ; University of Regensburg(莱茵河畔大学) ; Weill Cornell Medicine(韦尔·科恩医学中心) ; Northwell Health(北well健康)
专题命中 图文多模态 :cross-modal(title);multimodal(abstract);image-text(abstract);分类 cs.CV
机构 * MBZUAI ; Provable Responsible AI and Data Analytics (PRADA) Lab(可证明负责任的人工智能与数据分析实验室) ; King Abdullah University of Science and Technology(卡迪夫大学科学与技术大学) ; School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉科学学院) ; University of Copenhagen(哥本哈根大学) ; Tsinghua University(清华大学)
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV
专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI
专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV
Comments Accepted to NeurIPS 2025
机构 * University of Oxford, UK(牛津大学) ; CODE University of Applied Sciences, Germany(应用科学学院)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.AI
Journal ref Proceedings of the The European Workshop on Trustworthy AI (Trust-AI) at ECAI 2025
机构 * Shandong University(山东大学) ; MBZUAI ; New York University Abu Dhabi(纽约大学阿布扎赫分校) ; University of Surrey(塞雷尔大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
机构 * Simula Metropolitan Center for Digital Engineering (SimulaMet)(Simula数字工程研究中心) ; Oslo Metropolitan University (OsloMet)(奥斯陆 Metropolitan 大学) ; Simula Research Laboratory(Simula研究实验室)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
机构 * KingSoft Office Zhuiguang AI Lab(金山办公紫光人工智能实验室) ; Huazhong University of Science and Technology(华中科技大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI
机构 * Intellifusion Inc.(Intellifusion公司) ; Northwest Polytechnical University(西北工业大学) ; Harbin Institute of Technology(哈尔滨工业大学)
专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.AI
Comments 18 pages, 8 figures
专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.AI
Comments Accepted by AAAI2026
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV
Comments Accepted by AAAI 2026
机构 * New York University(纽约大学) ; Cornell Tech(康奈尔科技) ; Vanderbilt University(范德比尔特大学) ; Weill Cornell Medicine(韦尔医学院) ; Northwell Health(北well健康)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV
Journal ref Proceedings of SPIE Medical Imaging 2026
专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV
Comments AAAI2026
专题命中 图文多模态 :multimodal(abstract);分类 cs.CL
机构 * FusionBrain Lab(融合脑实验室) ; HSE University(俄罗斯高等经济学院) ; Innopolis University(因诺波利斯大学)
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
机构 * Monash University(墨尔本大学) ; Nanyang Technological University(南洋理工大学) ; The University of Hong Kong(香港大学) ; VinUni
专题命中 图文多模态 :multimodal(abstract);分类 cs.CV
Comments Updates to v1. Added new coauthors and extended the experimental section
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI、cs.MM
机构 * Google(谷歌)
专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.CV
Journal ref Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology (UIST '25), Article 17, 1-16, 2025
专题命中 音频语音多模态 :multimodal(title,abstract)
机构 * University of Pennsylvania(宾夕法尼亚大学)
专题命中 音频语音多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV、cs.CL
机构 * South China Normal University, Guangzhou, China(华南师范大学) ; Xiamen Rekey Medical Technology Co., LTD, Xiamen, China(厦门瑞康医疗科技有限公司)
专题命中 音频语音多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI
Comments 8 pages, 2 figures. To appear in: Proceedings of the 28th European Conference on Artificial Intelligence (ECAI 2025), Frontiers in Artificial Intelligence and Applications, Vol. 413. DOI: 10.3233/FAIA251182
机构 * University of Tsukuba(茨口大学) ; Cluster Metaverse Lab(元宇宙集群实验室) ; The University of Tokyo(东京大学) ; ZOZO Research(ZOZO研究所)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CV、cs.MM
Comments 3 pages, 2 figures, 1 table. Presented at SIGGRAPH Asia 2025 Posters (SA Posters '25), December 15-18, 2025, Hong Kong, Hong Kong
机构 * The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ; Prometheus Vision Technology Co., Ltd.(普罗米修斯视觉科技有限公司) ; The Hong Kong University of Science and Technology(香港科学与技术大学)
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV
Comments Proceedings of the Computer Graphics International 2025 (CGI'25)
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CL
Comments This article will serve as an extension of the preceding work, "VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models" (arXiv:2505.15727). Therefore, we have chosen to withdraw to avoid potential duplicate publication. We will update the previously open-sourced paper of VocalBench in several weeks to include the content of VocalBench-zh
专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV
Comments Our paper has been accepted to IEEE TPAMI
专题命中 视频多模态 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI
Comments AAAI 2026 accepted
专题命中 视频多模态 :multi-modal(title,abstract);cross-modal(abstract)
Comments 11 pages, 6 figures,