发表机构
Fudan University; Shanghai Innovation Institute; Ant Group(复旦大学; 上海创新研究院; 蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VLZip 是统一视觉与文本压缩的框架,通过提炼软前缀缩短注意力序列,在长上下文多模态推理上性能领先,支持超大规模 token 处理,为长上下文多模态 AI 建立新高效标准。
AI 中文摘要
视觉语言模型(VLMs)在处理超长、交错的图文序列时面临重大挑战,原因在于自注意力机制的二次复杂度。现有解决方案要么采用激进的 token 剪枝,存在不可逆信息丢失的风险;要么采用高效但精度较低的架构,且大多忽视了同样重要的文本组件。我们提出 VLZip,这是一个在纯 Transformer 框架内实现高保真推理的统一视觉与文本压缩框架。VLZip 的核心是将视觉和文本片段分层提炼为紧凑的、层特定的“软前缀”,并将其注入每个解码器层的隐藏状态,从而大幅缩短注意力序列,同时保留细粒度的全局上下文。为解决该领域评估不足的问题,我们还提出了 LongVLBench,一个源自视频叙事的新基准,要求进行整体的叙事级推理。大量实验表明,VLZip 在长上下文多模态推理上达到领先性能,支持最多 120K token 的训练,较基线提升 6 倍,推理可超过 280K token 且内存显著降低,同时具备处理多达 2M token 的内存可扩展性。在现有方法失效的极端上下文长度场景中表现优异,VLZip 为长上下文多模态 AI 建立了高效且强大的新标准。代码可在该 https URL 获取。
英文摘要
Vision Language Models (VLMs) face significant challenges with ultra-long, interleaved image-text sequences due to the quadratic complexity of self-attention. Current solutions either resort to aggressive token pruning, risking irreversible information loss, or adopt efficient but less precise architectures, while largely ignoring the equally vital textual component. We introduce VLZip, a framework that unifies visual and textual compression for high-fidelity reasoning within a pure Transformer. At its core, VLZip hierarchically distills visual and textual segments into compact, layer-specific "soft prefixes" and injects them into each decoder layer's hidden states, drastically shortening the attention sequence while preserving fine-grained global context. To address deficient evaluations in the field, we also introduce LongVLBench, a new benchmark derived from video narratives that demands holistic, narrative-level reasoning. Extensive experiments show VLZip achieves leading performance on long-context multimodal reasoning, enabling training up to 120K tokens, a 6x increase over the baseline, and inference beyond 280K tokens with significantly reduced memory, while demonstrating the memory scalability to handle up to 2M tokens. By excelling at extreme context lengths where existing methods collapse, VLZip establishes an efficient and powerful new standard for long-context multimodal AI. Code is available at https://github.com/ShareLab-SII/VLZip.