发表机构
Xiaomi Corporation(小米公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态大模型输入分辨率静态导致的效率瓶颈,提出任务条件分辨率路由(TCRR),用轻量跨模态路由器动态决定压缩级别,在Qwen3-VL-8B上减少40.9%视觉FLOPs和53.7%延迟,保持性能。
AI 中文摘要
多模态大语言模型(MLLMs)的推理效率受到高分辨率输入所产生的大量视觉标记序列的严重制约,其计算成本呈二次方增长。现有方法主要聚焦于下游标记压缩,却忽视了一个根本性的上游低效问题:输入分辨率被当作一个静态的、与任务无关的超参数。我们提出了任务条件分辨率路由(TCRR),将视觉压缩形式化为一个任务条件下的决策,并采用一个轻量级跨模态路由器,通过特征级调制和交叉注意力,将骨干视觉表示与文本语义进行条件化,以预测每个查询所需的最小充分压缩级别。为支持这一方法,我们构建了一个包含12个任务类别、50万个样本的数据集,通过教师-预言机流程进行标注,以近似帕累托最优的压缩尺度。跨多种架构的大量实验表明,TCRR实现了优越的效率前沿,具体而言,在Qwen3-VL-8B上将视觉FLOPs减少了40.9%,延迟降低了53.7%,同时保持了具有竞争力的性能。对扩展行为的进一步分析证实,动态路由视觉压缩能够在不修改MLLM骨干网络的情况下实现最优资源分配。
英文摘要
The inference efficiency of Multimodal Large Language Models (MLLMs) is severely constrained by massive visual token sequences induced by high-resolution inputs, with computational cost scaling quadratically. Existing approaches primarily focus on downstream token compression, while overlooking a fundamental upstream inefficiency: input resolution is treated as a static, task-agnostic hyperparameter. We propose Task-Conditioned Resolution Routing (TCRR), which formulates visual compression as a task-conditioned decision and employs a lightweight cross-modal router that conditions backbone visual representations on textual semantics via feature-wise modulation and cross-attention to predict the minimal sufficient compression level per query. To support this, we curate a dataset of 500k samples across 12 task categories, labeled via a teacher-oracle pipeline to approximate Pareto-optimal compression scales. Extensive experiments across diverse architectures show that TCRR achieves a superior efficiency frontier, specifically reducing visual FLOPs by 40.9% and latency by 53.7% on Qwen3-VL-8B while preserving competitive performance. Further analysis of scaling behavior confirms that dynamically routing visual compression enables optimal resource allocation without modifying the MLLM backbone.
Comments21 pages including references and appendix