LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs
LLaVA-UHD v3:渐进式视觉压缩用于多模态大语言模型中的高效原分辨率编码
Shichu Sun, Yichen Zhang, Haolin Song, Zonghao Guo, Chi Chen, Yidan Zhang, Yuan Yao, Zhiyuan Liu, Maosong Sun
机构
*
Tsinghua University(清华大学)
;
University of Chinese Academy of Sciences(中国科学院大学)
;
Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
;
Aerospace Information Research Institute, Chinese Academy of Sciences(中国科学院航天信息研究所)
;
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉学科学院)
机构
*
School of Cyber Science and Technology, University of Science and Technology of China(中国科学技术大学网络安全学院)
;
Anhui Province Key Laboratory of Digital Security(安徽省数字安全重点实验室)
;
The University of Hong Kong(香港大学)