Grounding Multilingual Multimodal LLMs With Cultural Knowledge
机构 * Carnegie Mellon University(卡内基梅隆大学)
专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
机构 * Carnegie Mellon University(卡内基梅隆大学)
专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CL
机构 * West Virginia University(西弗吉尼亚大学) ; University of Aberdeen(阿伯丁大学)
专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.CV
Comments 3rd Workshop in Data Engineering in Medical Imaging (DEMI), MICCAI-2025 Workshop
机构 * Fraunhofer IGD and Department of Computer Science, TU Darmstadt(弗劳恩霍夫研究所(IGD)和图宾根大学计算机科学系)
专题命中 图文多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.CV
Comments Accepted at ACM Multimedia Workshops
机构 * Peking University(北京大学) ; Institute of Artificial Intelligence (TeleAI), China Telecom(人工智能研究所(TeleAI),中国电信) ; Northwestern Polytechnical University(西北工业大学) ; Southeast University(东南大学)
专题命中 图文多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.AI
专题命中 图文多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; University of Washington(华盛顿大学) ; Shenzhen Institutes of Advanced Technology(深圳先进技术研究所) ; Chinese Academy of Sciences(中国科学院) ; State Key Laboratory of Ophthalmology(眼科学国家重点实验室) ; Zhongshan Ophthalmic Center(中山眼科中心) ; Sun Yat-sen University(中山大学) ; Guangdong Provincial Key Laboratory of Ophthalmology and Visual Science(广东省眼科学与视觉科学重点实验室) ; Guangdong Provincial Clinical Research Center for Ocular Diseases(广东省眼科临床研究中心) ; Department of Bioengineering(生物工程系) ; Department of Radiology(放射科)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV、cs.CL、cs.AI
Comments Accepted by IEEE TPAMI, 14 pages, 15 tables, 4 figures with Appendix
机构 * University of Chinese Academy of Sciences(中国科学院大学) ; Dept. of Comp. Sci. and Tech., Institute for AI, Tsinghua University(计算机科学与技术系,人工智能研究院,清华大学) ; Gaoling School of Artificial Intelligence, Renmin University of China(人工智能学院,中国人民大学) ; Kuaishou Technology Inc.(快手科技有限公司) ; Shanghai Key Laboratory of Multi. Info. Processing, East China Normal University(多信息处理重点实验室,华东师范大学)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV、cs.CL
机构 * Indian Institute of Science(印度科学研究院)
专题命中 图文多模态 :image-text(abstract);分类 cs.CV
专题命中 音频语音多模态 :MLLM(title,abstract);multimodal(abstract);分类 cs.MM、eess.AS
机构 * Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences, China(中国科学院自动化研究所基础模型研究中心)
专题命中 音频语音多模态 :audio-visual(title,abstract);cross-modal(abstract);分类 cs.CV
Comments Accpted by ICCV 2025
机构 * Faculty of Engineering and Architecture(工程与建筑学院) ; IDLab-AIRO, Ghent University – imec(IDLab-AIRO,根特大学–imec) ; Department of Engineering Science, University of Oxford(工程科学系,牛津大学)
专题命中 音频语音多模态 :multimodal(title,abstract)
专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.MM、eess.AS
Comments 17 pages, 12 figures, source code available at https://github.com/wguo86/SSV2A
机构 * Pennsylvania State University(宾夕法尼亚州立大学) ; Tsinghua University(清华大学)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI
Comments 10 pages, accepted to MRAC'25: 3rd International Workshop on Multimodal and Responsible Affective Computing (ACM-MM 2025)
机构 * School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院) ; Ministry of Education Key Laboratory of Intelligent Networks and Network Security, Xi’an Jiaotong University(西安交通大学教育部长江网络与网络安全重点实验室) ; College of Computer and Mathematics, Central South University of Forestry and Technology(中南林业科技大学计算机与数学学院) ; Department of Computer Science, State University of New York(纽约州立大学新帕尔茨分校计算机科学系)
专题命中 视频多模态 :multi-modal(title);multimodal(abstract);分类 cs.CL
Comments 13 pages, 6 figures
专题命中 视频多模态 :multimodal(title,abstract)
Comments 17 pages, 9 figures
机构 * School of Data Science, Fudan University(复旦大学数据科学学院)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
专题命中 视频多模态 :cross-modal(abstract);分类 cs.CV
机构 * University of Science, VNU-HCM(越南胡志明市科学大学) ; University of Information Technology, VNU-HCM(越南胡志明市信息技术大学) ; Vietnam National University(越南国家大学) ; University of Dayton(戴维森大学)
专题命中 跨模态检索 :multi-modal(title);分类 cs.CV
机构 * Munich University of Applied Sciences(慕尼黑应用科学大学) ; Intelligent Vehicles Lab (IVL)(智能车辆实验室)
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV、cs.CL
Comments Project page: https://iv.ee.hm.edu/contextmotionclip/; This work has been submitted to the IEEE for possible publication
专题命中 多模态生成 :multimodal(title);multi-modal(abstract);分类 cs.CV
机构 * organization= Center for Geospatial Sciences, Applications ; Department of Landscape of Architecture ; Urban Planning, Texas A\&M University , addressline= 788 Ross St , city= College Station , postcode= 77840 , state= TX , country= United States ; organization= Zachry Department of Civil ; Environmental Engineering, Texas A\&M University , addressline= 201 Dwight Look Engineering Building , city= College Station , postcode= 77843 , state= TX , country= United States ; organization= Department of Civil ; Environmental Engineering, University of Wisconsin-Madison , addressline= 1415 Engineering Dr , city= Madison , postcode= 53706 , state= WI , country= United States
专题命中 多模态生成 :multi-modal(title,abstract);分类 cs.AI
专题命中 多模态生成 :multimodal(title,abstract)
Comments 17 pages; 1 table; 6 figures; extended version of accepted version, published at the 2025 Winter Simulation Conference (WSC '25)
机构 * The University of Tokyo(东京大学) ; CyberAgent AI Lab(CyberAgent AI实验室)
专题命中 多模态生成 :multimodal(title);分类 cs.CV
Comments Accepted to ICDAR2025
机构 * CMU(卡内基梅隆大学) ; WUSTL(华盛顿大学) ; UIUC(伊利诺伊大学香槟分校)
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
机构 * School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院) ; School of Artificial Intelligence, Sun Yat-sen University(中山大学人工智能学院)
专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments Accepted by TPAMI
专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV
Comments This paper has been accepted by ICCV 2025 Workshop MMFM4
专题命中 多模态生成 :multimodal(abstract)
Comments 9 pages, 6 figures
机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) ; The University of Hong Kong(香港大学) ; Shanghai Jiao Tong University(上海交通大学) ; The Hong Kong University of Science and Technology(香港科学与技术大学) ; Fudan University(复旦大学) ; Tsinghua University(清华大学)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
机构 * Shenzhen University(深圳大学) ; University of Nottingham Ningbo China(诺丁汉大学宁波分校) ; City University of Hong Kong(香港城市大学) ; Stanford University(斯坦福大学) ; Hong Kong University of Science and Technology(香港科学与技术大学)
专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.CL
Comments ICCV 2025, 38 pages, 22 figures, 35 tables
机构 * Tel-Aviv University(特拉维夫大学)
专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV
Comments Accepted to ACL 2024 (Finding). For Project webpage, see https://moranyanuka.github.io/icc/
Journal ref Findings of the Association for Computational Linguistics: ACL 2024, pages 11048-11064, Bangkok, Thailand, August 2024