UltraViT:用于大型视觉语言模型的延迟优化设备端视觉编码器
UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
查看机构详情
- Samsung AI Cambridge(三星人工智能剑桥实验室)
- Technical University of Iasi(雅西技术大学)
- Queen Mary University of London(伦敦玛丽女王大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对大型视觉语言模型在设备端部署受限问题,提出UltraViT视觉编码器,通过考虑设备延迟设计金字塔架构,并采用两阶段生成式预训练策略,有效提升编码效率,优于现有基线。
中文摘要 AI 辅助
大型视觉语言模型(LVLMs)因巨大的计算量而受限,无法在资源有限的边缘设备上部署。以往压缩LVLMs的努力主要集中在减少视觉令牌或使用更小的语言模型,视觉编码器常被忽视。本文提出UltraViT,专为设备端性能设计优化。考虑实际设备延迟,系统设计金字塔架构,在宏块级别整合异构空间混合器。还提出两阶段生成式预训练策略,实验表明该策略能有效实现高级语义基础,结合设备端延迟设计的方法显著优于现有基线,实现高效LVLM编码。
英文摘要
Large Vision-Language Models (LVLMs) remain bottlenecked by massive computational footprints, precluding their deployment on resource-constrained edge devices. While efforts to compress LVLMs focus heavily on vision token reduction or smaller language models, the vision encoder is largely overlooked, typically deployed as a monolithic, computationally heavy feature extractor. Moreover, there is no previous effort that designs a vision encoder for LVLMs directly optimized for on-device latency. In this paper, we present UltraViT, a vision encoder for LVLMs, explicitly designed and optimized for on-device performance. Specifically, by taking into account real on-device latencies, we systematically design a pyramidal architecture that strategically integrates and adapts heterogeneous spatial mixers at the macro-block level. Furthermore, to pre-train UltraViT, we propose a novel two-stage generative pre-training strategy: cultivating rich spatial features via dense distillation, followed by direct generative supervision from a capacity-mixed frozen LLM. Compared to standard contrastive and SSL, we show that our pre-training is much more effective for achieving high-level semantic grounding for UltraViT needed for the subsequent generative multimodal alignment of LVLM training. Extensive experiments demonstrate that our on-device latency-informed design combined with our tailored training strategy establishes a new state-of-the-art for efficient LVLM encoding, significantly outperforming existing encoder-centric baselines while operating on-device at nearly 1.7xthe speed.