AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity
机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) ; Pattern Recognition Center, WeChat AI, Tencent Inc, China(腾讯人工智能研究院) ; Key Laboratory of Digital Protection and Intelligent Processing of Intangible Cultural Heritage of Fujian and Taiwan (Xiamen University), Ministry of Culture and Tourism, China(福建省和台湾非物质文化遗产数字化保护与智能处理重点实验室) ; Shanghai Artificial Intelligence Laboratory, China(上海人工智能实验室)
专题命中 VLM训练与架构 :LLaVA(title,abstract);分类 cs.CV、cs.AI
Comments Accepted by ACL 2025 Findings