arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

仅用MLP:多模态语言模型的高效视觉状态重建

Just MLPs: Efficient Visual State Reconstruction for Multimodal Language Models

Jingdi lei, Junxian Li, Di Zhang, Zhanqiu Zhang, Yiwen Guo, Soujanya Poria

arXiv 2609.34972首次发表:更新:

发表机构

Nanyang Technological University; Shanghai Jiao Tong University; Fudan University; LIGHTSPEED(南洋理工大学; 上海交通大学; 复旦大学; LIGHTSPEED)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态大语言模型中视觉token序列计算开销大的问题,本文提出δ-Vision方法,用轻量级低秩适配器替代Transformer对视觉token的反复演化,保留全部视觉token,在图像和视频基准上以相当或更低计算量取得更高准确率。

AI 中文摘要

在多模态大语言模型(MLLMs)中,长视觉token序列通常占计算开销的很大一部分。现有方法通过剪枝冗余视觉token来降低这一成本,但会永久丢弃可能在后续层中有用的视觉证据。我们转而探究是否可以在降低通过Transformer反复演化表示的代价的同时保留所有视觉token。为回答此问题,我们对视觉到文本的信息流进行低秩干预。我们发现,在阻断视觉到文本的注意力后,仅恢复少数方向即可恢复大部分丢失的准确率,这表明相关的视觉影响集中在低维子空间中。我们进一步观察到层特定视觉状态具有强可预测性:轻量级MLP以高余弦相似度和低重建误差近似它们。受这些发现启发,我们提出δ-Vision,该方法用轻量级低秩适配器替换视觉token的重复Transformer演化,构建逐层视觉记忆,同时保留所有视觉token供文本检索。在图像和视频基准上,δ-Vision在相当或更低计算量下比视觉token剪枝基线获得更高准确率,同时在不丢弃视觉token的情况下提供有竞争力的推理效率。

英文摘要

Long visual token sequences often account for a substantial fraction of the computational overhead in multimodal large language models~(MLLMs). Existing approaches reduce this cost by pruning redundant visual tokens, but permanently discard visual evidence that may become useful in subsequent layers. We instead ask whether all visual tokens can be preserved while reducing the cost of repeatedly evolving the representations through the Transformer. To answer this question, we perform low-rank interventions on visual-to-text information flow. We find that, after visual-to-text attention is blocked, restoring only a few directions recovers most of the lost accuracy, suggesting the relevant visual influence is concentrated in a low-dimensional subspace. We further observe strong predictability in layer-specific visual states: lightweight MLPs approximate them with high cosine similarity and low reconstruction error. Motivated by these findings, we propose $δ$-Vision, which replaces repeated Transformer evolution of visual tokens with lightweight low-rank adapters that construct layer-wise visual memories while preserving all visual tokens for text retrieval. Across image and video benchmarks, $δ$-Vision achieves higher accuracy than visual token pruning baselines at comparable or lower computation, while delivering competitive inference efficiency without discarding visual tokens.

Comments21 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑