CommentsPlease cite the definitive, copyrighted, and peer-reviewed version of this article published in AAAI 2026, edited by Sven Koenig et al., AAAI Press, Vol. 40, No. 36, Technical Track, pp. 30726-30734, 2026. DOI: https://doi.org/10.1609/aaai.v40i36.40329
Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach
重新审视多模态KV缓存压缩:一种基于频域的异常KV感知方法
Yaoxin Yang, Peng Ye, Xudong Tan, Chongjun Tu, Maosen Zhao, Jia Hao, Tao Chen
机构
*
College of Future Information Technology, Fudan University(未来信息科技学院,复旦大学)
;
Shanghai Innovation Institute(上海创新研究院)
;
The Chinese University of Hong Kong(香港中文大学)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Zhangjiang Laboratory(张江实验室)
Induced Numerical Instability: Hidden Costs in Multimodal Large Language Models
诱导的数值不稳定性:多模态大语言模型中的隐藏成本
Wai Tuck Wong, Jun Sun, Arunesh Sinha
机构
*
School of Computing(计算学院)
;
Information Systems, Singapore Management University, Singapore(信息系统,新加坡管理大学,新加坡)
;
Information Systems Department, Rutgers Business School, New Jersey, USA(信息系统系,罗格斯商学院,新泽西,美国)
Has Multimodal Learning Delivered Universal Intelligence in Healthcare? A Comprehensive Survey
多模态学习是否在医疗领域实现了通用智能?一项全面的综述
Qika Lin, Yifan Zhu, Xin Mei, Ling Huang, Jingying Ma, Kai He, Zhen Peng, Erik Cambria, Mengling Feng
机构
*
Saw Swee Hock School of Public Health, National University of Singapore(新加坡国立大学公共健康学院)
;
School of Computer Science, Beijing University of Posts and Telecommunications(北京邮电大学计算机学院)
;
School of Automation, Northwestern Polytechnical University(西北工业大学自动化学院)
;
School of Computer Science and Technology, Xi’an Jiaotong University(西安交通大学计算机科学与技术学院)
Deep Multimodal Learning with Missing Modality: A Survey
缺失模态下的深度多模态学习:综述
Renjie Wu, Hu Wang, Hsiang-Ting Chen, Gustavo Carneiro
机构
*
The Australian National University(澳大利亚国立大学)
;
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
;
Adelaide University(阿德莱德大学)
;
The University of Surrey(萨里大学)
机构
*
Faculty of Computer Science, AGH University of Krakow(计算机科学系,克拉科夫AGH大学)
;
Department of Informatics, Universitas Pembangunan Nasional Veteran Yogyakarta(信息系,全国 veterans 大学 Yogya 市)
ADMN: A Layer-Wise Adaptive Multimodal Network for Dynamic Input Noise and Compute Resources
Jason Wu, Yuyang Yuan, Kang Yang, Lance Kaplan, Mani Srivastava
机构
*
Electrical and Computer Engineering University of California, Los Angeles(电气与计算机工程大学加州大学洛杉矶分校)
;
DEVCOM Army Research Laboratory(国防部陆军研究实验室)
;
University of California, Los Angeles(加州大学洛杉矶分校)
;
Amazon(亚马逊)
Macro2Micro: A Rapid and Precise Cross-modal Magnetic Resonance Imaging Synthesis using Multi-scale Structural Brain Similarity
Sooyoung Kim, Joonwoo Kwon, Junbeom Kwon, Jungyoun Janice Min, Sangyoon Bae, Yuewei Lin, Shinjae Yoo, Jiook Cha
机构
*
Department of Brain and Cognitive Science, Seoul National University, Seoul, Republic of Korea(脑科学与认知科学系,首尔国立大学)
;
Department of Applied Bioengineering, Seoul National University, Seoul, Republic of Korea(应用生物工程系,首尔国立大学)
;
Department of Psychology, Seoul National University, Seoul, Republic of Korea(心理学系,首尔国立大学)
;
Interdisciplinary Program in Artificial Intelligence, Seoul National University, Seoul, Republic of Korea(人工智能跨学科项目,首尔国立大学)
;
Computational Science Initiative, Brookhaven National Laboratory, Upton, NY, USA(计算科学计划,布鲁赫斯国家实验室)
;
School of Economics, Sogang University(经济学院,成均馆大学)
;
Brookhaven National Laboratory(布鲁赫斯国家实验室)