机构
*
School of Computer Science and Engineering, Northeastern University, Shenyang, China(东北大学计算机科学与工程学院)
;
Key Laboratory of Intelligent Computing in Medical Image of Ministry of Education, Northeastern University, Shenyang, China(教育部医学图像智能计算重点实验室)
;
National Frontiers Science Center for Industrial Intelligence and Systems Optimization, Shenyang, China(工业智能与系统优化国家级前沿科学中心)
;
Alberta Machine Intelligence Institute, University of Alberta, Edmonton, Canada(阿尔伯塔机器智能研究所,阿尔伯塔大学,加拿大爱德蒙顿)
机构
*
Zhejiang Key Laboratory of Space Information Sensing and Transmission, Hangzhou Dianzi University(浙江空间信息感知与传输重点实验室,杭州电子大学)
;
Laboratory of Complex Systems Modeling and Simulation, School of Computer Science and Technology, Hangzhou Dianzi University(复杂系统建模与仿真实验室,计算机科学与技术学院,杭州电子大学)
;
Alibaba International Digital Commerce Group(阿里巴巴国际数字商业集团)
;
College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)
Disentangled Multi-modal Learning of Histology and Transcriptomics for Cancer Characterization
解耦的多模态学习:组织学与转录组学用于癌症表征
Yupei Zhang, Xiaofei Wang, Anran Liu, Lequan Yu, Chao Li
机构
*
Department of Clinical Neurosciences, University of Cambridge, UK(剑桥大学临床神经科学系)
;
Department of Health Technology & Informatics, The Hong Kong Polytechnic University(香港理工大学健康科技与信息学系)
;
Department of Statistics and Actuarial Science, The University of Hong Kong(香港大学统计与精算科学系)
;
Department of Clinical Neurosciences and Department of Applied Mathematics and Theoretical Physics, University of Cambridge(剑桥大学临床神经科学系和应用数学与理论物理系;邓迪大学科学与工程学院和医学院)
;
School of Science and Engineering and School of Medicine, University of Dundee, UK
Not All Attention is Needed: Parameter and Computation Efficient Transfer Learning for Multi-modal Large Language Models
并非所有注意力都是必需的:面向多模态大语言模型的参数和计算高效迁移学习
Qiong Wu, Weihao Ye, Yiyi Zhou, Xiaoshuai Sun, Rongrong Ji
机构
*
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China(教育部多媒体可信感知与高效计算重点实验室)
;
Institute of Artificial Intelligence, Xiamen University(厦门大学人工智能研究院)
Trifuse: Enhancing Attention-Based GUI Grounding via Multimodal Fusion
Trifuse: 通过多模态融合增强基于注意力的GUI定位
Longhui Ma, Di Zhao, Siwei Wang, Zhao Lv, Miao Wang
机构
*
College of Computer Science and Technology, National University of Defense Technology(计算机科学与技术学院,国防科技大学)
;
Intelligent Game and Decision Lab, Academy of Military Sciences(智能游戏与决策实验室,军事科学院)
The Paradigm Shift: A Comprehensive Survey on Large Vision Language Models for Multimodal Fake News Detection
范式转变:大型视觉语言模型在多模态虚假新闻检测中的全面调查
Wei Ai, Yilong Tan, Yuntao Shou, Tao Meng, Haowen Chen, Zhixiong He, Keqin Li
机构
*
College of Computer and Mathematics, Central South University of Forestry and Technology(计算机与数学学院,中央南大学林业科技学院)
;
College of Computer Science and Electronic Engineering, Hunan University(计算机科学与电子工程学院,湖南大学)
;
College of Economics and Management, Central South University of Forestry and Technology(经济管理学院,中央南大学林业科技学院)
;
Department of Computer Science, State University of New York(计算机科学系,纽约州立大学)
机构
*
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉学科学院)
;
State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统国家重点实验室)
;
Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science and Innovation, Chinese Academy of Sciences(中国科学院香港创新科学研究院人工智能与机器人中心)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Fujian Key Laboratory of Pattern Recognition and Image Understanding, School of Computer and Information Engineering, Xiamen University of Technology(福建 pattern recognition and image understanding 工程学院,厦门大学科技学院)
Multimodal Learning for Scalable Representation of High-Dimensional Medical Data
多模态学习用于高维医学数据的可扩展表示
Areej Alsaafin, Abubakr Shafique, Saghir Alfasly, Krishna R. Kalari, H. R. Tizhoosh
机构
*
Kimia Lab, Dept. of Artificial Intelligence & Informatics, Mayo Clinic, Rochester, MN, USA(Kimia实验室,人工智能与信息学系,梅奥诊所,罗切斯特,MN,美国)
;
Division of Computational Biology, Dept. of Quantitative Health Sciences, Mayo Clinic, Rochester, MN, USA(计算生物学部门,定量健康科学系,梅奥诊所,罗切斯特,MN,美国)
机构
*
School of Information Science & Engineering, Shandong Normal University(信息科学与工程学院,山东师范大学)
;
Tajikistan State University of Law, Business Sughd(塔吉克斯坦法律、商业大学,苏赫德)
;
Tajik State University of Law, Business and Politics Sughd(塔吉克斯坦法律、商业与政治大学,苏赫德)
Unlabeled Data Improves Fine-Grained Image Zero-shot Classification with Multimodal LLMs
未标记数据提升多模态大语言模型在细粒度图像零样本分类中的性能
Yunqi Hong, Sohyun An, Andrew Bai, Neil Y. C. Lin, Cho-Jui Hsieh
机构
*
Computer Science Department, University of California, Los Angeles(加州大学洛杉矶分校计算机科学系)
;
Mechanical and Aerospace Engineering Department, University of California, Los Angeles(加州大学洛杉矶分校机械与航空航天工程系)