PACE:用于快速视觉语言模型推理的统一压缩-提取范式
PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference
浏览论文内容
中文总结 AI 辅助
针对视觉语言模型推理成本随视觉 token 增多而上升的问题,提出无需训练的 PACE 框架,集成到 Qwen2.5-VL-7B 后仅用 10% 视觉 token 保留 93.8% 原性能,TTFT 加速 3.1 倍。
中文摘要 AI 辅助
视觉语言模型(Vision-Language Models, VLMs)展现出卓越的视觉推理能力,但其推理成本会随视觉 token 的增多而快速上升。现有的视觉 token 剪枝方法存在两个核心局限:其一,多数方法仅在视觉编码器之后运行,未优化视觉编码阶段的大量延迟;其二,在严格的 token 预算下,这些方法往往无法同时保留整体视觉上下文与细粒度细节,导致性能下降。为解决这些瓶颈,我们提出 PACE(Pixel-Adaptive Condense and Extract,像素自适应压缩与提取),这是一种无需训练的推理框架,通过统一的压缩-提取范式同时加速视觉编码器与大语言模型(Large Language Model, LLM)。在压缩阶段,自适应像素压缩器(Adaptive Pixel Compressor, APC)在编码前评估视觉信息密度,自适应下采样冗余输入,在保留全局上下文与关键视觉线索的同时减少编码器计算量;在提取阶段,动态双注意力提取器(Dynamic Dual-Attention Extractor, DDAE)通过融合编码器的内部视觉信号与 LLM 的语义信号,选择性保留视觉 token,保障任务关键细节。将 PACE 集成到 Qwen2.5-VL-7B 后,模型仅使用 10% 的视觉 token,同时保留了原模型 93.8% 的性能,首 token 生成时间(Time to First Token, TTFT)实现了 3.1 倍加速。我们的代码可在该 https URL 获取。
英文摘要
Vision-Language Models (VLMs) demonstrate exceptional visual reasoning capabilities, yet their inference costs escalate rapidly with the proliferation of visual tokens. Existing visual token pruning methods exhibit two fundamental limitations. First, most approaches operate exclusively post-vision encoder, leaving the substantial latency of the visual encoding phase unoptimized. Second, under strict token budgets, these methods often fail to jointly preserve holistic visual contexts and fine-grained details, leading to performance degradation. To address these bottlenecks, we propose PACE (Pixel-Adaptive Condense and Extract), a training-free inference framework that accelerates both the vision encoder and the Large Language Model (LLM) via a unified Condense-and-Extract paradigm. During the Condense stage, an Adaptive Pixel Compressor (APC) evaluates visual information density prior to encoding, adaptively downsampling redundant inputs, curtailing encoder computation while preserving global context and essential visual cues. In the Extract stage, a Dynamic Dual-Attention Extractor (DDAE) selectively retains visual tokens via a fusion of internal visual signals from the encoder and semantic signals from the LLM, safeguarding task-critical details. By integrating PACE into Qwen2.5-VL-7B, the model retains 93.8% of its original performance while utilizing only 10% of the visual tokens, yielding a 3.1x speedup in time to first token (TTFT). Our code is available at https://github.com/jjL357/PACE.
发表机构
- Sun Yat-sen University(中山大学)
- Power Dispatch Control Center, Guangdong Power Grid Co., Ltd.(广东电网有限责任公司电力调度控制中心)
- Shenzhen Loop Area Institute(深圳河套学院)
机构由 AI 辅助整理,请以论文原文为准。