arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01614cs.CVcs.AI

线性多时间尺度保留作为内存高效的视觉-语言桥梁

Linear Multi-Timescale Retention as a Memory-Efficient Vision-Language Bridge

Ashfak Yeafi, Mehedi Hasan, Md Khairul Islam

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对视觉-语言模型高分辨率图像处理的内存瓶颈,提出线性多时间尺度保留模块,兼具O(N)复杂度与上下文路由能力,在长序列任务、硬件效率及MME基准上均优于MLP基线。

中文摘要 AI 辅助

视觉-语言模型(VLMs)在处理高分辨率图像时面临严重的计算瓶颈,原因是Softmax多头注意力(MHA)的内存复杂度为O(N²)。虽然用独立的多层感知机(MLP)替代MHA可实现O(N)的缩放,但这会使架构失去空间序列路由能力,严重降低全局场景理解和目标持久性。本文提出线性多时间尺度保留(LIA-MTR)模块,一种内存高效的跨模态桥梁。通过集成基于ELU的正特征映射、自适应写门控和对数线性分布的循环衰减,LIA-MTR从数学上将连续视觉序列压缩为有界记忆状态。理论分析证明该架构具有严格的O(N)序列交互复杂度。实验方面,合成检索评估显示LIA-MTR可完美在16000个token间路由上下文,消除了朴素线性注意力典型的“中间丢失”退化;硬件基准测试显示其具备无限上下文缩放能力,可在11.2GB显存内原生处理262144个视觉patch,而标准MHA在16384个patch时就会出现内存不足。此外,在665K个对话样本上进行指令调优后,LIA-MTR在MME基准上显著优于行业标准MLP基线(71.00% vs. 68.11%),这源于其在目标持久性上10%的绝对提升和更优的全局语义提取。本工作为无限上下文视觉-语言集成建立了数学严谨、计算平稳的基础。

英文摘要

Vision-Language Models (VLMs) face a critical computational bottleneck when processing high-resolution imagery due to the $O(N^2)$ memory complexity of Softmax Multi-Head Attention (MHA). While substituting MHA with independent Multi-Layer Perceptrons (MLPs) achieves $O(N)$ scaling, it strips the architecture of spatial sequence routing, severely degrading global scene understanding and object permanence. In this paper, we propose the Linear Multi-Timescale Retention (LIA-MTR) module, a memory-efficient cross-modal bridge. By integrating an ELU-based positive feature mapping with adaptive write-gating and log-linearly distributed recurrent decays, LIA-MTR mathematically compresses continuous visual sequences into bounded memory states. Theoretical analysis proves the architecture operates with strict $O(N)$ sequence-interaction complexity. Empirically, synthetic retrieval evaluations demonstrate that LIA-MTR flawlessly routes context across 16,000 tokens, eliminating the "Lost in the Middle" degradation typical of naive linear attention. Hardware benchmarking reveals infinite-context scaling capabilities, natively processing 262,144 visual patches within an 11.2 GB VRAM footprint, whereas standard MHA suffers out-of-memory failure at 16,384 patches. Furthermore, following instruction tuning on 665K conversational samples, LIA-MTR significantly outperforms an industry-standard MLP baseline on the MME benchmark (71.00% vs. 68.11%), driven by a 10% absolute improvement in object permanence and superior global semantic extraction. This work establishes a mathematically rigorous, computationally flat foundation for infinite-context Vision-Language integration.

↑