AI 中文总结
该研究提出无训练框架LAST,利用边缘侧VLM的最后查询令牌注意力实现查询感知视觉令牌剪枝,在11个多模态基准上以12.5%视觉令牌保留率维持95.4%全令牌准确率,降低边缘和云侧开销。
AI 中文摘要
多模态基础模型正将边缘-云视觉智能从特定任务特征流水线重塑为基于令牌的接口,边缘设备将视觉输入编码为令牌后传输至通用云侧多模态大语言模型(MLLM)。然而,密集的视觉令牌序列会增加云侧推理成本。现有剪枝方法主要针对集中式推理:视觉驱动方法可在云执行前操作,但通常与查询无关;查询引导方法常依赖目标MLLM的内部状态,无法在传输前确定令牌相关性;紧凑引导模型是替代方案,但现有设计可能需要代价高昂的注意力聚合或辅助生成。我们提出LAST,一种用于边缘-云协作MLLM推理中查询依赖型视觉令牌剪枝的无训练框架。LAST使用紧凑的边缘侧视觉语言模型(VLM)作为引导代理,从最后一个查询令牌对视觉令牌的注意力中推导轻量级重要性信号。在因果注意力机制下,最后一个查询令牌可关注完整视觉序列和整个查询上下文,无需访问云模型、自回归生成或对多个查询位置进行代价高昂的聚合,即可实现查询感知剪枝。LAST在固定令牌预算下保留一组多样化的查询相关视觉令牌。我们在11个多模态基准上,针对不同引导策略的剪枝方法,在多种令牌预算下对LAST进行评估。实验表明,LAST始终取得最强性能,在仅保留12.5%视觉令牌的同时,保留了全令牌准确率的95.4%,且边缘侧选择开销低,云侧计算量减少。
英文摘要
Multimodal foundation models are reshaping edge-cloud visual intelligence from task-specific feature pipelines into token-based interfaces, where edge devices encode visual inputs into tokens for a general-purpose cloud MLLM. However, dense visual-token sequences increase cloud-side inference costs. Existing pruning methods mainly target centralized inference: vision-driven methods can operate before cloud execution but are typically query-agnostic, whereas query-guided methods often rely on internal states of the target MLLM and cannot determine token relevance before transmission. Compact guidance models offer an alternative, but existing designs may require costly attention aggregation or auxiliary generation. We propose LAST, a training-free framework for query-dependent visual token pruning in edge-cloud collaborative MLLM inference. LAST uses a compact edge-side VLM as a guidance proxy and derives a lightweight importance signal from the last query token's attention to visual tokens. Under causal attention, the last query token can attend to the full visual sequence and the entire query context, enabling query-aware pruning without cloud-model access, autoregressive generation, or costly aggregation over multiple query positions. LAST then retains a diverse set of query-relevant visual tokens under a fixed token budget. We evaluate LAST on 11 multimodal benchmarks under multiple token budgets against pruning methods with different guidance strategies. Experiments show that LAST consistently achieves the strongest performance, preserving 95.4% of the full-token accuracy while retaining only 12.5% of the visual tokens, with low edge-side selection overhead and reduced cloud-side computation.