HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
HiRes-LLaVA: 高分辨率大视觉-语言模型中碎片化输入的恢复
机构 * Shenzhen campus of Sun Yat-sen University(中山大学深圳校区) ; Huawei(华为) ; The Hong Kong University of Science and Technology(香港科技大学) ; The University of Hong Kong(香港大学)
专题命中 VLM训练与架构 :vision-language model(title,abstract);LLaVA(title,abstract);分类 cs.CV
AI总结 HiRes-LLaVA通过引入SliceRestore适配器和Self-Mining Sampler,有效恢复高分辨率输入的碎片化问题,提升模型在文档相关任务中的性能。