arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LookME:用于视觉语言模型层注入的基于查找的多模态嵌入

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models

Zeyu Xu, Xingzhong Hou, Pengkai Guo, Siling Lin, Xiao Xu, Menghua Zhai, Haoyu Chen, Yunke Zhang, Fei Huang

arXiv 2607.16305首次发表:更新:

AI 中文总结

本文针对视觉语言模型在资源受限环境中部署的问题,提出LookME框架。通过分层两级查找方法和稀疏注入策略,实现基于查找的多模态嵌入增强,实验表明该方法优于仅文本的PLE风格方法,有效提升了模型性能。

AI 中文摘要

视觉语言模型在多模态理解方面取得了显著进展。然而,扩展密集或稀疏专家混合模型以提高性能,由于全加载的高内存使用和按需加载的延迟增加之间的权衡,限制了在资源受限环境中的部署。最近,逐层嵌入(PLE)架构通过使用存储在ROM中的大型外部嵌入表扩展模型并执行轻量级查找来检索相关嵌入以增强令牌表示来解决此问题。然而,现有的PLE风格方法主要针对文本嵌入设计,限制了其在视觉语言模型中的有效性。本文提出了LookME,第一个支持基于查找的多模态嵌入增强的框架,同时支持分区存储和按需加载。为了从大规模嵌入表中高效查找任意连续多模态嵌入,提出了一种分层两级查找方法,采用从场景级到场景内原语级的粗到细策略。此外,将查找方法与稀疏注入策略集成,自适应地优先考虑关键嵌入,促进跨相邻层的嵌入表重用,改善效率、模型大小和性能之间的权衡。在多个视觉基准上的实验表明,LookME优于仅文本的PLE风格方法,验证了基于查找的多模态嵌入增强的有效性。

英文摘要

Vision-Language Models (VLMs) have achieved strong progress in multimodal understanding. However, scaling dense or sparse Mixture-of-Experts (MoE) models to improve performance limits deployment in resource-constrained environments due to the trade-off between high memory usage from full loading and increased latency from on-demand loading. Recently, the Per-Layer Embedding (PLE) architecture addresses this by scaling models with large external embedding tables stored in read-only memory (ROM) and performing lightweight lookup to retrieve relevant embeddings to enhance token representations. Nevertheless, existing PLE-style methods are primarily designed for text embeddings due to the convenience of ID-based retrieval, limiting their effectiveness in VLMs where multimodal embeddings contain richer information for visual tasks. In this paper, we propose LookME, the first framework that enables lookup-based enhancement for multimodal embeddings in VLMs while supporting partitioned storage and on-demand loading. To efficiently lookup arbitrary continuous multimodal embeddings from large-scale embedding tables, we propose a hierarchical two-level lookup method employing a coarse-to-fine strategy that performs lookups from the scene-level to the intra-scene primitive-level. Furthermore, we integrate the lookup method with a sparse injection strategy, which adaptively prioritizes critical embeddings over voluminous multimodal embeddings within layers, and facilitates embedding table reuse across neighboring layers, improving the trade-off among efficiency, model size, and performance. Experiments on multiple visual benchmarks show that LookME outperforms text-only PLE-style methods, validating the effectiveness of lookup-based multimodal embedding enhancement.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑