发表机构
Purdue University; University of Wisconsin–Madison; NVIDIA(普渡大学; 威斯康星大学麦迪逊分校; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态大语言模型推理成本高的问题,提出SmartVL框架,通过视觉侧令牌控制器和LLM侧计算控制器联合控制视觉令牌数量和模型计算能力,实验证明该框架优于先前方法,实现更好的精度-效率平衡。
AI 中文摘要
多模态大语言模型(MLLMs)在视觉-语言任务中表现出色,但推理成本高阻碍实际部署。近期工作尝试单独优化各维度来降低成本,却忽视了计算资源需根据输入内容动态分配这一耦合关系。为此提出SmartVL统一自适应推理框架,它通过视觉侧令牌控制器和LLM侧计算控制器联合控制视觉令牌数量和模型计算能力,并使其相互协调以满足目标预算。实验表明,SmartVL优于先前自适应方法,实现了更好的精度-效率帕累托前沿。
英文摘要
Multimodal Large Language Models (MLLMs) have recently demonstrated strong performance across vision-language tasks. However, their high inference cost, arising from both the large number of input visual tokens and the heavy computation of the large language model (LLM), remains a key barrier to practical deployment. Recent work attempts to reduce the cost by adaptively optimizing individual dimensions, e.g., pruning redundant visual tokens or skipping LLM layers and heads. Nonetheless, prior approaches typically treat these dimensions independently and overlook a fundamental coupling: the available compute resources must be dynamically allocated across all dimensions based on the input content. To bridge the gap, we propose SmartVL, a unified adaptive inference framework that jointly controls vision token number and model compute capability in response to varying input contents and compute budgets. SmartVL introduces a vision-side token controller that dynamically selects informative visual tokens and an LLM-side compute controller that adaptively adjusts LLM computation. Importantly, these controllers are trained to coordinate with each other so that the overall inference cost satisfies a target budget. To allow this joint scheduling, we connect the controllers using a shared budget encoding and leverage a differentiable latency estimator for end-to-end training. This design enables SmartVL to learn cross-stage allocation strategies that adapt to both input complexity and runtime compute constraints. Experiments across multiple MLLM benchmarks demonstrate that, with joint scheduling, SmartVL consistently outperforms prior adaptive methods and achieves superior accuracy-efficiency Pareto frontiers. Project page: https://www.schaterji.io/publications/2026/jointtokencompute.
CommentsAccepted at ECCV 2026