arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26596cs.CVcs.AI

解耦视觉处理:通过模态特定Transformer替换实现高效多模态适配

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution

  • Tsinghua University(清华大学)
  • Beijing National Research Center for Information Science and Technology(北京信息科学与技术国家研究中心)

机构由 AI 辅助整理,请以论文原文为准。

Mingkuan Feng, Zhengqi Wen, Jianhua Tao

AI总结:

该研究针对多模态大语言模型视觉指令微调成本高的问题,提出解耦视觉处理框架,替换预训练大语言模型上层解码器为轻量Transformer块,仅训练该块即可在多个基准上实现竞争力性能,降低了计算成本。

AI中文摘要:

多模态大语言模型(MLLMs)通过在统一Transformer架构中整合视觉与文本理解,展现出卓越能力。然而,对这类模型的所有参数进行视觉指令微调计算成本高昂且常无必要,因为网络深层中视觉与文本token的表征需求存在显著差异。本文提出解耦视觉处理(Decoupled Visual Processing, DVP),这是一种高效训练框架,将预训练大语言模型(LLM)的上层解码器层替换为轻量、可独立训练的单一Transformer块,专门用于视觉token处理。具体而言,经过解码器前半部分的共享处理后,视觉与文本token被拆分:视觉token被路由至新初始化的单一Transformer块,而文本token则继续通过原始冻结的解码器层。两条流在语言建模头前被拼接。训练期间,仅更新该单一Transformer块,大幅减少可训练参数数量。在LLaVA-1.5框架上的实验表明,DVP在MME、POPE和ChartQA基准上实现了具有竞争力的性能,同时仅训练总参数的一小部分,这表明MLLMs中的视觉表征可通过解耦的参数高效路径有效学习。

英文摘要:

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual understanding within a unified transformer architecture. However, fine-tuning all parameters of these models for visual instruction tuning is computationally expensive and often unnecessary, as the representation requirements for visual and textual tokens diverge significantly in the deeper layers of the network. In this paper, we propose Decoupled Visual Processing (DVP), an efficient training framework that replaces the upper decoder layers of a pretrained LLM with a lightweight, independently trainable single transformer block dedicated exclusively to visual token processing. Specifically, after shared processing through the first half of the decoder layers, visual and textual tokens are split: visual tokens are routed through a newly initialized single transformer block while textual tokens continue through the original frozen decoder layers. The two streams are then concatenated before the language modeling head. During training, only the single transformer block is updated, dramatically reducing the number of trainable parameters. Experiments on the LLaVA-1.5 framework demonstrate that DVP achieves competitive performance on MME, POPE, and ChartQA benchmarks while training only a fraction of the total parameters, suggesting that visual representations in MLLMs can be effectively learned through a decoupled, parameter-efficient pathway.

↑