Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models
潜在噪声掩码:减少多模态大语言模型中的视觉冗余
Kai Jiang, Ruishu Zhu, Siqi Huang, Hongyuan Zhang, Xuelong Li
机构
*
School of Artificial Intelligence, OPtics and ElectroNics (iOPEN), Northwestern Polytechnical University(人工智能学院、光学和电子学(iOPEN)、西北工业大学)
;
Institute of Artificial Intelligence, China Telecom (TeleAI)(人工智能研究院、中国电信(TeleAI))
;
Fudan University(复旦大学)
;
The University of Hong Kong(香港大学)
TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference
TOPS:通过构建令牌最优保留集实现高效多模态大语言模型推理的第一性原理视觉令牌剪枝
Tinghao Wang, Yichen Guo, Rui Huang, Zheng Lu, Qizhe Zhang, Chenxi Li, Yuan Zhang, Jiajun Cao, Zhirong Shen, Yaosong Du, Guangyan Gan, Wenya Wang, Lin William Cong, Shanghang Zhang
机构
*
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(北京大学计算机学院多媒体信息处理国家重点实验室)
;
University of Electronic Science and Technology of China(电子科技大学)
;
Nanyang Technological University(南洋理工大学)
;
Beijing Academy of Artificial Intelligence (BAAI)(北京人工智能研究院)
Spectral Evolution-Guided Token Pruning in Multimodal Large Language Models
多模态大语言模型中基于谱演化引导的令牌剪枝
Bin Chen, Yuxiang Cai, Yadan Luo, Yi Zhang, Jianwei Yin, Zhi Chen
机构
*
School of Software Technology, Zhejiang University(浙江大学软件学院)
;
Zhejiang Key Laboratory of Digital-Intelligence Service Technology(浙江省数字化服务技术重点实验室)
;
The University of Queensland(昆士兰大学)
;
Singapore Management University(新加坡管理大学)
;
The University of Southern Queensland(南昆士兰大学)
Cross-Modal Knowledge Distillation without Paired Data: Theoretical Foundation and Algorithm
无配对数据的跨模态知识蒸馏:理论基础与算法
Trong Khiem Tran, Anh Duc Chu, Quang Hung Pham, Phi Le Nguyen, Trong Nghia Hoang
机构
*
School of Information and Communications Technology, Hanoi University of Science and Technology, Hanoi, Vietnam(信息与通信技术学院,河内科学技术大学,越南河内)
;
School of Electrical Engineering and Computer Science, Washington State University, Pullman, US(电气工程与计算机科学学院,华盛顿州立大学,华盛顿州普尔曼)
CommentsThis article presents only the preliminary research results, which are not yet complete and lack necessary supplementary experiments. The author has decided to withdraw it to improve the research work, and will submit a more complete version in the future
MultiPUFFIN: A Multimodal Domain-Constrained Foundation Model for Molecular Property Prediction of Small Molecules
MultiPUFFIN:用于小分子性质预测的多模态领域约束基础模型
Idelfonso B. R. Nogueira, Carine M. Rebello, Mumin Enis Leblebici, Erick Giovani Sperandio Nascimento
机构
*
Department of Chemical Engineering, Norwegian University of Science and Technology (NTNU)(挪威科学与技术大学化学工程系)
;
Faculty of Industrial Engineering, KU Leuven(鲁文大学工业工程学院)
;
University of Surrey(萨里大学)
专题命中
多模态训练与对齐
:multimodal(title,abstract);cross-modal(abstract);multimodal foundation model(abstract);分类 cs.AI
机构
*
School of Software Engineering, the National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, National Engineering Research Center for Visual Information and Applications, and Institute of Artificial Intelligence and Robotics(软件工程学院、人机混合增强智能国家重点实验室、视觉信息与应用国家工程研究中心、人工智能与机器人研究院)
;
School of Software Engineering(软件工程学院)
;
XJTU-POLIMI Joint School(西交大-波兰理工联合学院)
;
Faculty of Electronic and Information Engineering(电子与信息工程学院)
;
School of Human Settlements and Civil Engineering(人居与土木工程学院)
;
the National Key Laboratory of Human-Machine Hybrid Augmented Intelligence, National Engineering Research Center for Visual Information and Applications, and Institute of Artificial Intelligence and Robotics(人机混合增强智能国家重点实验室、视觉信息与应用国家工程研究中心、人工智能与机器人研究院)
Better with Less: Tackling Heterogeneous Multi-Modal Image Joint Pretraining via Conditioned and Degraded Masked Autoencoder
更少的协同,更好的表现:通过条件和降质掩码自编码器解决异质多模态图像联合预训练
Bowen Peng, Yongxiang Liu, Jie Zhou, Xiaodong Chen, Tianpeng Liu, Xiaogang Yu, Li Liu
机构
*
College of Electronic Science and Technology, National University of Defense Technology (NUDT)(电子科学与技术学院,国防科技大学)
;
Beijing Institute of Remote Sensing Information(遥感信息研究所)
机构
*
MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所多模态人工智能系统实验室)
;
King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学)
;
Pengcheng Laboratory(鹏城实验室)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS)(中国科学院大学人工智能学院)
CoVFT: Context-aware Visual Fine-tuning for Multimodal Large Language Models
CoVFT:面向多模态大语言模型的上下文感知视觉微调
Nan Zhou, Huiqun Wang, Yaoyan Zheng, Di Huang
机构
*
State Key Laboratory of Complex and Critical Software Environment, Beihang University(北京航空航天大学复杂关键软件环境国家重点实验室)
;
School of Computer Science and Engineering, Beihang University(北京航空航天大学计算机科学与工程学院)
机构
*
Pengcheng Laboratory(鹏城实验室)
;
Shaanxi University of Science & Technology(陕西科技大学)
;
School of Computer Science(计算机科学学院)
;
Northwestern Polytechnical University(西北工业大学)
SEF-MAP: Subspace-Decomposed Expert Fusion for Robust Multimodal HD Map Prediction
SEF-MAP:子空间分解专家融合用于鲁棒多模态高精度地图预测
Haoxiang Fu, Lingfeng Zhang, Hao Li, Ruibing Hu, Zhengrong Li, Guanjing Liu, Zimu Tan, Long Chen, Hangjun Ye, Xiaoshuai Hao
机构
*
National University of Singapore(新加坡国立大学)
;
Xiaomi EV(小米电动车)
;
Chinese University of Hong Kong(香港中文大学)
;
The University of Manchester(曼彻斯特大学)
;
Renmin University of China(中国人民大学)
Multi-Modal Sensing and Fusion in mmWave Beamforming for Connected Vehicles: A Transformer Based Framework
毫米波波束成形中多模态感知与融合:基于变换器的框架
Muhammad Baqer Mollah, Honggang Wang, Mohammad Ataul Karim, Hua Fang
机构
*
Department of Electrical and Computer Engineering, University of Massachusetts Dartmouth(电子与计算机工程系,马萨诸塞大学达特茅斯分校)
;
Department of Graduate Computer Science and Engineering, Katz School of Science and Health, Yeshiva University(研究生计算机科学与工程系,耶鲁大学科学与健康学院)
Beyond Next-Token Alignment: Distilling Multimodal Large Language Models via Token Interactions
超越单个词对齐:通过令牌交互蒸馏多模态大语言模型
Lin Chen, Xiaoke Zhao, Kun Ding, Weiwei Feng, Changtao Miao, Zili Wang, Wenxuan Guo, Ying Wang, Kaiyuan Zheng, Bo Zhang, Zhe Li, Shiming Xiang
机构
*
MAIS, Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所信息与智能系统研究所)
;
School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院)
;
Zhejiang University(浙江大学)