arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ParVL:面向多模态大语言模型的并行缩放与可扩展计算分配

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs

Yang Yang, Qinyu Zhao, Mouxiang Chen, Xiaohui Li, Lixin Gu, Wenhai Wang, Hongjie Zhang, Wenwei Zhang

arXiv 2608.04010首次发表:更新:

AI 中文总结

本研究针对多模态大语言模型的计算分配僵化问题,提出ParVL框架,通过复用骨干参数扩展并行计算,在13B token上微调后,其多模态性能优于同配置单分支基线,且最优分配随任务变化。

AI 中文摘要

现有多模态大语言模型(MLLM)的缩放策略通常要么扩展模型参数,要么扩展顺序推理计算,这会带来大量内存或延迟开销。更重要的是,大多数现有方法无法改变视觉Transformer(ViT)与大语言模型(LLM)组件之间僵化、固定的计算分配,限制了针对特定任务的优化。为解决这一问题,我们提出面向MLLM的并行视觉-语言(ParVL)缩放框架,该框架通过在多个视觉和语言分支中复用现有ViT与LLM骨干参数来扩展并行计算。该框架提出一个核心问题:在固定的骨干参数预算下,应如何在视觉和语言模态之间分配额外的共享骨干计算?我们在共享骨干上为每个并行计算流实例化分支特定的前缀参数,并通过在约130亿个token上进行全参数监督微调来端到端训练整个模型。我们系统研究了ViT编码器与LLM解码器之间的计算分配权衡。ParVL相比同配置的单分支基线提升了整体多模态性能,且评估得到的最优视觉-语言分配随任务不同而变化。代码可在this https URL获取。

英文摘要

Existing scaling strategies for Multimodal Large Language Models (MLLMs) typically expand either model parameters or sequential inference computation, incurring substantial memory or latency overhead. More importantly, most existing methods fail to alter the rigid, fixed computation allocation between the Vision Transformer and the Large Language Model components, limiting task-specific optimization. To address this, we introduce the Parallel Vision-Language (ParVL) scaling framework for MLLMs, which scales parallel computation by reusing the existing ViT and LLM backbone parameters across multiple vision and language branches. This framework raises a central question: given a fixed backbone parameter budget, how should additional shared-backbone computation be allocated between the vision and language modalities? We instantiate each parallel computational stream with branch-specific prefix parameters over a shared backbone, and train the entire model end-to-end via full-parameter supervised fine-tuning on roughly 13B tokens. We systematically study the computation-allocation trade-off between the ViT encoder and LLM decoder. ParVL improves overall multimodal performance over same-recipe single-branch baselines, and the best evaluated vision--language allocation varies across tasks. Code is available at https://github.com/YangYangGirl/ParVL.

Comments14 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑