发表机构
Anyscale; Mississippi State University; Institute of Computing Technology, Chinese Academy of Sciences; Institute of AI for Industries, Chinese Academy of Sciences(安尼斯科尔公司; 密西西比州立大学; 中国科学院计算技术研究所; 中国科学院产业人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态大语言模型推理开销大的问题,提出具备自适应边云协作的Pro-Router方法,通过两阶段渐进式决策机制提升路由准确率与速度,端到端吞吐量显著优于现有方案。
AI 中文摘要
多模态大语言模型(Multimodal Large Language Models, MLLMs)虽性能卓越,但计算开销巨大,给实时部署与成本效益带来重大挑战。现有模型路由方法要么仅从粗粒度的请求级特征做决策,要么需花费一次或多次额外的大语言模型前向传播来检查生成的响应,未利用生成过程中出现的令牌级不确定性信号。为解决这些局限,本文提出Pro-Router,一种面向高效多模态大语言模型推理的、具备自适应边云协作的令牌感知渐进式模型路由方法。Pro-Router采用两阶段渐进式决策机制:第一阶段,轻量级提示预评分器模块在令牌生成开始前执行快速预筛选,将明显简单的请求导向小型模型;第二阶段,令牌感知验证器读取小型模型生成的每个令牌的采样概率分布,评估模型对自身输出的置信度,以逐请求决定是交付答案还是将请求升级至基于云端的高精度模型。此外,本文设计了自适应边云服务流水线,根据各设备实测服务速率调整每次调度的规模,使边缘与云端层级均保持充分利用,无需人工参数调优,且不受网络延迟影响。在多个多模态基准数据集与模型上开展的大量实验验证了Pro-Router的有效性:与其他方法相比,它实现了最高的路由准确率,路由速度提升超10倍;其服务流水线的端到端吞吐量比现有模型路由流水线高出75%以上。本文代码可在该https URL获取。
英文摘要
The remarkable performance of multimodal large language models (MLLMs) comes at the cost of substantial computational overhead, posing significant challenges to real-time deployment and cost effectiveness. Existing model routing approaches either decide from coarse request-level features alone or spend one or several extra language model passes to inspect the generated response, leaving the token-level uncertainty signals that emerge during generation unused. To address these limitations, we propose Pro-Router, a token-aware progressive model routing method with adaptive edge-cloud collaboration for efficient multimodal LLM inference. Pro-Router employs a two-stage progressive decision mechanism. First, a lightweight prompt pre-scorer module performs rapid pre-screening before token generation begins, guiding apparently simple requests to small models. Second, a token-aware verifier reads the sampling probability distribution of each token the small model generates, estimating the model's confidence in its own output to determine, per request, whether the answer ships or escalates to the cloud-based high-precision model. Furthermore, we design an adaptive edge-cloud serving pipeline that sizes every dispatch to each device's measured service rate, so both the edge and the cloud tiers stay fully utilized without manual parameter tuning and are not impacted by the network latency. Extensive experiments on multiple multimodal benchmark datasets and models demonstrate the effectiveness of Pro-Router. Compared to other methods, it achieves the highest routing accuracy and improves routing speed by more than 10x. Its serving pipeline also reaches more than 75% higher end-to-end throughput than the existing model routing pipeline. Our code is available at https://github.com/xinyuangui2/pro-router.
Comments9 pages, 7 figures, 2 tables. Code: https://github.com/xinyuangui2/pro-router