arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于交通管理的成本最优基础模型部署组合

Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management

Xi Cheng, Ke Liu, Siyuan Feng, Jane Lin, H. Oliver Gao

arXiv 2607.13239首次发表:更新:

发表机构

Cornell University; University of California, Berkeley; The Hong Kong Polytechnic University; University of Illinois Chicago(康奈尔大学; 加州大学伯克利分校; 香港理工大学; 伊利诺伊大学芝加哥分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究交通管理中基础模型部署组合问题,提出FMDP混合整数规划,证明其NP难,给出多项式时间贪婪启发式算法,通过案例研究确定低成本组合,经盈亏平衡分析得出本地GPU投资合理的条件。

AI 中文摘要

基础模型,包括大语言模型(LLMs)和视觉语言模型(VLMs),越来越多地用于交通管理中心(TMC)任务,如异常检测、事件报告和旅行者信息。在TMC功能中部署多个此类模型会引发一个组合问题:每个功能应由哪个模型以何种部署模式在何种共享硬件预算下提供服务?我们将此表述为基础模型部署组合(FMDP)问题,这是一个混合整数规划问题,在共享GPU容量上,使总拥有成本(TCO)最小化,同时满足每个功能的质量、延迟和安全约束。我们通过从0-1背包问题归约证明该问题为NP难问题,并提出一种多项式时间贪婪启发式算法。在一个具有五个TMC功能和19个候选(模型,模式)对的案例研究中,FMDP通过将四个功能路由到开源API,将一个没有开源模型能满足其质量底线的功能路由到封闭API,确定了一个每月成本为34美元的混合组合(比最便宜的可行全封闭API基线低97%)。盈亏平衡分析表明,只有当每小时视觉查询量约高于309次或API价格翻倍时,本地GPU投资才变得合理。

英文摘要

Foundation models, including large language models (LLMs) and vision-language models (VLMs), are increasingly used for transportation management center (TMC) tasks such as anomaly detection, incident reporting, and traveler information. Deploying multiple such models across TMC functions raises a portfolio question: which model should serve each function, in which deployment mode, and under what shared hardware budget? We formulate this as the Foundation Model Deployment Portfolio (FMDP) problem, a mixed-integer program minimizing total cost of ownership (TCO) subject to per-function quality, latency, and safety constraints over shared GPU capacity. We prove the problem NP-hard by reduction from the 0-1 knapsack problem and propose a polynomial-time greedy heuristic. In an illustrative case study with five TMC functions and 19 candidate (model, mode) pairs, FMDP identifies a mixed portfolio costing $34/mo (97% below the cheapest feasible all-closed-API baseline) by routing four functions to open-source APIs and the one function whose quality floor no open-source model meets to a closed API. Break-even analysis shows that on-premise GPU investment becomes reasonable only above approximately 309 vision queries/hour or if API prices double.

CommentsAccepted at IEEE ITSC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑