LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture
机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; Lehigh University(莱斯利大学) ; Meituan(美团)
专题命中 视频理解 :video understanding(abstract);分类 cs.CV、cs.MM
Comments Accepted to EMNLP 2025 Findings