Understanding vs. Generation: Navigating Optimization Dilemma in Multimodal Models
理解与生成:多模态模型中的优化困境导航
Sen Ye, Mengde Xu, Shuyang Gu, Di He, Liwei Wang, Han Hu
机构
*
State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
;
Tencent(腾讯)
;
Center for Data Science, Peking University(北京大学数据科学中心)
;
Center for Machine Learning Research, Peking University(北京大学机器学习研究中心)
Reflect to Inform: Boosting Multimodal Reasoning via Information-Gain-Driven Verification
反思以获取信息:通过信息增益驱动的验证提升多模态推理
Shuai Lv, Chang Liu, Feng Tang, Yujie Yuan, Aojun Zhou, Kui Zhang, Xi Yang, Yangqiu Song
机构
*
University of Science and Technology of China(中国科学技术大学)
;
Huawei Foundation Model Department(华为基础模型部)
;
The Chinese University of Hong Kong(香港中文大学)
;
Hong Kong University of Science and Technology(香港科技大学)
机构
*
College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机科学与人工智能学院)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
;
School of Artificial Intelligence, Shanghai Jiao Tong University(上海交通大学人工智能学院)
Multiscale Structure-Guided Latent Diffusion for Multimodal MRI Translation
多尺度结构引导的潜在扩散模型用于多模态MRI翻译
Jianqiang Lin, Zhiqiang Shen, Peng Cao, Jinzhu Yang, Osmar R. Zaiane, Xiaoli Liu
机构
*
Computer Science and Engineering, Northeastern University, Shenyang, China(东北大学计算机科学与工程系,沈阳,中国)
;
Key Laboratory of Intelligent Computing in Medical Image of Ministry of Education, Northeastern University, Shenyang, China(教育部医学图像智能计算重点实验室,东北大学,沈阳,中国)
;
National Frontiers Science Center for Industrial Intelligence and Systems Optimization, Shenyang, China(工业智能与系统优化国家级前沿科学中心,沈阳,中国)
;
Amii, University of Alberta, Edmonton, Alberta, Canada(阿尔伯塔大学阿米人工智能研究所,埃德蒙顿,阿尔伯塔,加拿大)
;
AiShiWeiLai AI Research, Beijing, China(人工智能研究(北京)有限公司,北京,中国)
Efficient endometrial carcinoma screening via cross-modal synthesis and gradient distillation
通过跨模态合成与梯度蒸馏实现高效的子宫内膜癌筛查
Dongjing Shan, Yamei Luo, Jiqing Xuan, Lu Huang, Jin Li, Mengchu Yang, Zeyu Chen, Fajin Lv, Yong Tang, Chunxiang Zhang
机构
*
School of Medical Information and Engineering, Southwest Medical University(西南医科大学医学信息与工程学院)
;
School of Computer Science and Engineering, University of Electronic Science and Technology of China(电子科技大学计算机科学与工程学院)
;
Department of Ultrasound, Affiliated Hospital of Southwest Medical University(西南医科大学附属医院超声科)
;
Department of Functional Examination Unit, Zibo Hospital of Traditional Chinese Medicine(淄博中医药医院功能检查科)
;
Key Laboratory of Medical Electrophysiology, Ministry of Education& Medical Electrophysiological Key Laboratory of Sichuan Province, Institute of Cardiovascular Research, Southwest Medical University(教育部医学电生理重点实验室、四川省医学电生理重点实验室、西南医科大学心血管研究所)
;
Department of Radiology, the First Affiliated Hospital of Chongqing Medical University(重庆医科大学第一附属医院放射科)
;
Department of Cardiology, Affiliated Hospital of Southwest Medical University(西南医科大学附属医院心内科)
;
International Research Center for Complexity Sciences, Hangzhou International Innovation Institute, Beihang University(北京航空航天大学杭州国际创新研究院复杂科学研究中心)
;
Institute of Intelligent Chinese Medicine, Chongqing University of Chinese Medicine(重庆中医药大学智能中药研究院)
;
Basic Medicine Research Innovation Center for Cardiometabolic Diseases, Ministry of Education, Southwest Medical University(教育部心脑血管疾病基础医学研究创新中心、西南医科大学)
Training-Free Text-Guided Color Editing with Multi-Modal Diffusion Transformer
无需训练的文本引导颜色编辑与多模态扩散变换器
Zixin Yin, Xili Dai, Ling-Hao Chen, Deyu Zhou, Jianan Wang, Duomin Wang, Gang Yu, Lionel M. Ni, Lei Zhang, Heung-Yeung Shum
机构
*
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
International Digital Economy Academy(国际数字经济学院)
;
The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州))
;
Tsinghua University(清华大学)
;
Astribot
;
StepFun
SpecFLASH: A Latent-Guided Semi-autoregressive Speculative Decoding Framework for Efficient Multimodal Generation
SpecFLASH: 一种基于潜在引导的半自回归推测解码框架,用于高效多模态生成
Zihua Wang, Ruibo Li, Haozhe Du, Joey Tianyi Zhou, Yu Zhang, Xu Yang
机构
*
Southeast University, Nanjing, China(东南大学)
;
Nanyang Technological University, Singapore(南洋理工大学)
;
A STAR Centre for Frontier AI Research (CFAR), Singapore(A STAR前沿人工智能研究中心)
机构
*
Organizational Management Department, School of Management, Xi’an Jiaotong University(管理学院组织管理部,西安交通大学)
;
West China Longquan Hospital, Sichuan University(四川大学西部临床医学院)
;
School of Electronic Science and Engineering, Xi’an Jiaotong University(西安交通大学电子科学与工程学院)
;
Systems Engineering Institute, Xi’an Jiaotong University(西安交通大学系统工程研究院)
;
Institute of Medical Artificial Intelligence, the Second Affiliated Hospital of Xi’an Jiaotong University(西安交通大学第二附属医院医学人工智能研究所)
;
School of Human Settlements and Civil Engineering, Xi’an Jiaotong University(西安交通大学人居环境与土木工程学院)
;
School of Life Science and Technology, Xi’an Jiaotong University(西安交通大学生命科学与技术学院)
Generating crossmodal gene expression from cancer histopathology improves multimodal AI predictions
从癌症组织病理学生成跨模态基因表达以提高多模态AI预测
Samiran Dey, Christopher R. S. Banerji, Partha Basuchowdhuri, Sanjoy K. Saha, Deepak Parashar, Tapabrata Chakraborti
机构
*
School of Mathematical & Computational Sciences, Indian Association for the Cultivation of Science(数学与计算科学学院,印度科学培养协会)
;
The Alan Turing Institute(艾伦·图灵研究所)
;
Comprehensive Cancer Center, King’s College London(国王学院综合癌症中心)
;
Department of Computer Science and Engineering, Jadavpur University(计算机科学与工程系,贾瓦德pur大学)
;
MRC Biostatistics Unit, University of Cambridge(剑桥大学医学研究委员会生物统计学单位)
;
Department of Biostatistics, Bioinformatics and Biomathematics, Georgetown University(生物统计学、生物信息学与生物数学系,杰斐逊大学)
;
UCL Cancer Institute, Dept of Medical Physics & Biomedical Engineering, University College London(伦敦大学学院癌症研究所,医学物理与生物医学工程系)