arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.23373cs.CV

UltraViT:用于大型视觉语言模型的延迟优化设备端视觉编码器

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models

发表机构三星人工智能剑桥实验室 · 雅西技术大学 · 伦敦玛丽女王大学
查看机构详情
  • Samsung AI Cambridge(三星人工智能剑桥实验室)
  • Technical University of Iasi(雅西技术大学)
  • Queen Mary University of London(伦敦玛丽女王大学)

机构由 AI 辅助整理,请以论文原文为准。

Ioannis Maniadis Metaxas, Adrian Bulat, Alberto Baldrati, Anestis Zaganidis, Yassine Ouali, Hyeonuk Kim, Georgios Tzimiropoulos

首次发表
浏览论文内容

中文总结 AI 辅助

针对大型视觉语言模型在设备端部署受限问题,提出UltraViT视觉编码器,通过考虑设备延迟设计金字塔架构,并采用两阶段生成式预训练策略,有效提升编码效率,优于现有基线。

中文摘要 AI 辅助

大型视觉语言模型(LVLMs)因巨大的计算量而受限,无法在资源有限的边缘设备上部署。以往压缩LVLMs的努力主要集中在减少视觉令牌或使用更小的语言模型,视觉编码器常被忽视。本文提出UltraViT,专为设备端性能设计优化。考虑实际设备延迟,系统设计金字塔架构,在宏块级别整合异构空间混合器。还提出两阶段生成式预训练策略,实验表明该策略能有效实现高级语义基础,结合设备端延迟设计的方法显著优于现有基线,实现高效LVLM编码。

英文摘要

Large Vision-Language Models (LVLMs) remain bottlenecked by massive computational footprints, precluding their deployment on resource-constrained edge devices. While efforts to compress LVLMs focus heavily on vision token reduction or smaller language models, the vision encoder is largely overlooked, typically deployed as a monolithic, computationally heavy feature extractor. Moreover, there is no previous effort that designs a vision encoder for LVLMs directly optimized for on-device latency. In this paper, we present UltraViT, a vision encoder for LVLMs, explicitly designed and optimized for on-device performance. Specifically, by taking into account real on-device latencies, we systematically design a pyramidal architecture that strategically integrates and adapts heterogeneous spatial mixers at the macro-block level. Furthermore, to pre-train UltraViT, we propose a novel two-stage generative pre-training strategy: cultivating rich spatial features via dense distillation, followed by direct generative supervision from a capacity-mixed frozen LLM. Compared to standard contrastive and SSL, we show that our pre-training is much more effective for achieving high-level semantic grounding for UltraViT needed for the subsequent generative multimodal alignment of LVLM training. Extensive experiments demonstrate that our on-device latency-informed design combined with our tailored training strategy establishes a new state-of-the-art for efficient LVLM encoding, significantly outperforming existing encoder-centric baselines while operating on-device at nearly 1.7xthe speed.

补充信息

↑