CT-CLIP: A Multi-modal Fusion Framework for Robust Apple Leaf Disease Recognition in Complex Environments
专题命中 图文多模态 :multi-modal(title);multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 图文多模态 :multi-modal(title);multimodal(abstract);image-text(abstract);分类 cs.CV、cs.AI
机构 * School of Automation, Guangdong University of Technology(广东工业大学自动化学院) ; College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics(南京航空航天大学计算机科学与技术学院) ; State Key Laboratory of Traditional Chinese Medicine Syndrome, The Second Affiliated Hospital of Guangzhou University of Chinese Medicine, Guangdong Provincial Hospital of Chinese Medicine, Guangdong Provincial Academy of Chinese Medical Sciences(广东省中医药科学院中医证候重点实验室,广州中医药大学第二附属医院,广东省中医院,广东省中医药科学院) ; College of Computing and Data Science, Nanyang Technological University(南洋理工大学计算机与数据科学学院) ; Department of Electrical and Computer Engineering, Sungkyunkwan University(成均馆大学电子与计算机工程系)
专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI
机构 * Guangdong Institute of Intelligence Science and Technology(广东智能科学与技术研究院) ; ZJU-Angelalign R&D Center for Intelligence Healthcare(浙大天使align智能医疗研发中心) ; Centre for Frontier AI Research (CFAR)(前沿人工智能研究中心) ; Agency for Science, Technology and Research (A*STAR)(科技研究局) ; Institute of High Performance Computing (IHPC)(高性能计算研究所)
专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV
机构 * Dalian University of Technology(大连理工大学) ; University of Electronic Science and Technology of China(电子科技大学) ; Tsinghua Shenzhen International Graduate School(清华大学深圳国际graduate school)
专题命中 图文多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CV
Comments Accepted to NeurIPS 2025
机构 * Key Lab of Intell. Info. Process., Inst. of Comput. Tech., CAS(智能信息处理重点实验室,计算技术研究所,中国科学院) ; University of Chinese Academy of Sciences(中国科学院大学)
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV
Comments Accepted by NeurIPS 2025
专题命中 图文多模态 :multi-modal(abstract);分类 cs.CV
Comments Accepted into the ICAAI 2025 - The 9th International Conference on Advances in Artificial Intelligence
机构 * Scientific Computing and Imaging Institute(科学计算与成像研究所) ; Department of Electrical and Computer Engineering(电气与计算机工程系) ; University of Utah(犹他大学)
专题命中 音频语音多模态 :cross-modal(title);audio-visual(title);分类 cs.CV
机构 * Worcester Polytechnic Institute(沃斯特理工大学) ; Amazon AGI(亚马逊人工智能实验室)
专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract);分类 eess.AS
机构 * Nanyang Technological University, Singapore(南洋理工大学) ; Peking University, China(北京大学) ; Shenzhen University, China(深圳大学)
专题命中 音频语音多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV、cs.AI、cs.MM
Comments NeurIPS 2025, Spotlight
机构 * State Key Laboratory for Novel Software Technology, Nanjing University(南京大学新型软件技术国家重点实验室) ; School of Intelligence Science and Technology, Nanjing University(南京大学智能科学与技术学院) ; Tencent Youtu Lab(腾讯优图实验室) ; XMU(厦门大学) ; CASIA(中国科学院自动化研究所)
专题命中 音频语音多模态 :MLLM(abstract,comments);multimodal(abstract);分类 cs.CV、eess.AS
Comments NeurIPS 2025 Spotlight, Code 2.4K Stars: https://github.com/VITA-MLLM/VITA
机构 * ADAPT Centre, School of Engineering(ADAPT中心,工程学院)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL、eess.AS
Comments Accepted to INTERSPEECH 2025, 10.21437/Interspeech.2025-668
Journal ref Proc. of INTERSPEECH 2025, 1073--1077, 10.21437/Interspeech.2025-668
机构 * Department of Computer Science, University of Florida, Gainesville, FL, USA(佛罗里达大学计算机科学系) ; Department of Computer Science, Lebanese American University(黎巴嫩美国大学计算机科学系)
专题命中 音频语音多模态 :cross-modal(abstract);分类 cs.CL、eess.AS
Comments 7 pages, 1 Figure, 8 tables, Under review ICASSP 2026
机构 * Stability AI
专题命中 音频语音多模态 :multi-modal(abstract);分类 cs.CV
Comments Project Page: https://stability-ai.github.io/foleycontrol.github.io/
机构 * ADAPT Centre, School of Engineering, Trinity College Dublin(ADAPT中心、工程学院、都柏林信任学院)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.CL
Comments Accepted to ACL 2025, Findings of the Association for Computational Linguistics
Journal ref In Findings of the Association for Computational Linguistics: ACL 2025, pages 209--221, Vienna, Austria. Association for Computational Linguistics, 10.18653/v1/2025.findings-acl.12
机构 * Dept. of Computer Science and Technology(计算机科学与技术系) ; University of Alicante(阿尔瓦登特大学) ; Unit of Clinical Nursing Research(临床护理研究单位) ; Faculty of Health Sciences(健康科学学院)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI
Journal ref Applied Soft Computing, Vol. 184, 2025, Article 113787
机构 * Amazon(亚马逊) ; Siebel School of Computing and Data Science, University of Illinois Urbana-Champaign(计算机与数据科学学院,伊利诺伊大学厄巴纳-香槟分校)
专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI
Comments 49 pages
机构 * 2 Department of Electrical ; Computer Engineering University of Waterloo, Waterloo, ON, Canada N2L 3G1 Email ; 3 College of Computer ; Information Sciences Prince Sultan University Email
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.CL、cs.AI
Journal ref 2024 IEEE International Symposium on Medical Measurements and Applications (MeMeA)
机构 * MoE Key Laboratory of Brain-Machine Intelligence Technology, College of Artificial Intelligence, Nanjing University of Aeronautics and Astronautics(脑机智能技术MoE实验室,人工智能学院,南京航空航天大学) ; Nanjing University(南京大学) ; The Hong Kong Polytechnic University(香港理工大学)
专题命中 视频多模态 :multi-modal(title,abstract);分类 cs.CV
Comments Accepted to NeurIPS 2025 D&B Track
机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校)
专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV
Comments Accepted by NeurIPS 2025, Project website: https://vision.cs.utexas.edu/projects/SeeAoT
机构 * Department of Computer Sciences University of Wisconsin-Madison(计算机科学系威斯康星大学麦迪逊分校)
专题命中 视频多模态 :multimodal(title,abstract)
Comments 15 pages, 8 figures
专题命中 视频多模态 :multi-modal(title)
Comments Accepted to ASE'25
机构 * State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, China(网络与交换技术国家重点实验室,北京邮电大学)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV、cs.AI
机构 * National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, School of Computer Science(多媒体软件国家工程研究中心、人工智能研究院、计算机科学学院) ; Hubei Key Laboratory of Multimedia and Network Communication Engineering(多媒体与网络通信工程湖北省重点实验室) ; School of Mathematical Sciences, Peking University(北京大学数学科学学院) ; Tsinghua University(清华大学) ; School of Computer Science, Peking University(北京大学计算机科学学院) ; State Key Laboratory for Multimedia Information Processing, Peking University(多媒体信息处理国家重点实验室)
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
机构 * Faculty of Engineering and Natural Sciences (VPALab)(工程与自然科学学院(VPALab)) ; Sabanci University(萨班奇大学)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV
专题命中 视频多模态 :multimodal(abstract);分类 cs.CV
Comments NeurIPS 2025
机构 * Vanderbilt University(范德比大学)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.AI
Comments Proceeding of the 39th Conference on Neural Information Processing Systems (NeurIPS'25). Code would be available at https://github.com/scope-lab-vu/ESCORT
机构 * McGill University(麦吉尔大学) ; Mila ; Stanford University(斯坦福大学) ; York University(约克大学) ; Polytechnique Montréal(蒙特利尔理工学院)
专题命中 视频多模态 :multi-modal(abstract);分类 cs.CV
Comments Project page: https://mtesfaldet.net/genpt_projpage/
专题命中 视频多模态 :multimodal(abstract)
Comments NeurIPS 2025
专题命中 跨模态检索 :multimodal(abstract);分类 cs.CV
Comments 17 pages, 4 figures
机构 * Google DeepMind(谷歌DeepMind)
专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI
Journal ref International Conference on Machine Learning, 2025