arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MoR-MLLM:高效多模态大语言模型的递归混合

MoR-MLLM: Mixture of Recursions for Efficient Multimodal Large Language Models

Pengcheng Zheng, Chaoning Zhang, Jiaxin Yan, Sihan Cao, Jianwei Zhang, Xudong Wang, Jiaquan Zhang, Jewon Lee, Tae-Ho Kim, Yang Yang, Heng Tao Shen

arXiv 2610.08830首次发表:更新:

发表机构

University of Electronic Science and Technology of China; Kyung Hee University; Nota Inc.; Tongji University(电子科技大学; 庆熙大学; Nota公司; 同济大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MoR-MLLM,基于递归混合框架实现计算稀疏的多模态大语言模型,通过自适应逐令牌递归动态分配计算资源,降低训练内存和计算复杂度,同时保持高性能。

AI 中文摘要

多模态大语言模型(MLLMs)在视觉和语言任务中展现出了卓越的推理能力。然而,其巨大的计算和内存需求阻碍了实际部署。尽管最近的工作通过采用轻量级语言骨干网络来降低成本,但现有范式由于静态稀疏性和深度分配,仍然计算密集,无法适应每个令牌的语义复杂度。为此,我们提出了MoR-MLLM,一种基于近期递归混合(MoR)框架的计算稀疏型MLLM。MoR-MLLM引入了自适应的逐令牌递归,使模型能够动态调整其递归深度,为视觉或语言上具有挑战性的令牌分配更多计算资源,同时跳过简单令牌的冗余操作。为了在多模态环境中稳定递归稀疏性的训练,我们进一步设计了三阶段MoR-Tuning策略和熵正则化损失,以鼓励多样化的路由分布。大量实验表明,与近期先进的小型MLLMs相比,我们提出的MoR-MLLM能够大幅降低训练内存和计算复杂度,同时在各种视觉-语言任务上保持高性能。

英文摘要

Multimodal Large Language Models (MLLMs) have demonstrated remarkable reasoning capabilities across vision and language tasks. However, their massive computational and memory demands hinder real-world deployment. While recent efforts reduce costs by employing lightweight language backbones, existing paradigms remain computation-dense due to their static sparsity and depth allocation, which cannot adapt to the semantic complexity of each token. To this end, we propose MoR-MLLM, a computation-sparse MLLM based on the recent Mixture-of-Recursions (MoR) framework. MoR-MLLM introduces adaptive per-token recursion, allowing the model to dynamically adjust its recursive depth and allocate more computation to visually or linguistically challenging tokens while skipping redundant operations for simpler ones. To stabilize the training of recursive sparsity in multimodal settings, we further design a three-stage MoR-Tuning strategy and an entropy-regularized loss to encourage diverse routing distributions. Extensive experiments show that compared with recent advanced tiny MLLMs, our proposed MoR-MLLM can greatly reduce the training memory and computation complexity while retaining high performance on various vision-language tasks.

Comments13 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑