你不能从价格表中选择提供商:面向开放权重LLM推理的市场感知路由
You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference
另 1 家 · 查看机构详情
- University of Sydney(悉尼大学)
- Johns Hopkins University(约翰斯·霍普金斯大学)
- Stanford University(斯坦福大学)
- City University of Hong Kong(香港城市大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对开放权重LLM推理市场,提出市场感知路由问题,通过测量映射与FACET在线认证机制,实现同模型下选择质优价廉且健康的提供商,节省成本并避免退化端点。
中文摘要 AI 辅助
现有的LLM路由器使用静态的每模型成本在模型之间进行选择。我们表明,开放权重推理市场引入了第二个在很大程度上被忽视的决策轴:在选择了模型之后,客户端仍然必须选择哪个提供商来服务它。我们测量了[nummodels]个开放模型、竞争提供商、多种任务类型和三个测量波次的实时端点,发现提供商的选择无法从价格表中推断出来。同一模型在不同提供商之间的质量、延迟、可用性和价格可能差异很大;价格较高的提供商始终更快,但价格并不能可靠地预测质量或可用性;并且提供商的可行性是任务选择性的,一个部署在知识任务上几乎正常,但在多步推理上灾难性地退化。我们将同模型提供商选择表述为一个价格接受者的市场感知路由问题。一个简单的测量映射策略路由到既质量等价又健康的、最便宜的提供商,从而在避免退化端点的同时实现匹配质量的节省。由于映射会漂移,我们引入了FACET,一个在线提供商路由器,它认证每个(提供商×任务)的可行性方面,并在服务未认证的臂之前安全地回退到锚点。在放宽的部署假设下,FACET容忍不完美的任务分配和稀疏反馈,而系统性评估器偏差暴露了一个质量信号信任边界,可以通过地面真值探针或审计来缓解。实时提供商运行进一步证实,认证可以将实际流量从优质锚点转移到更便宜的认证端点。我们的结果表明,市场感知的LLM路由必须不仅测量使用哪个模型,还要测量谁服务它。
英文摘要
Existing LLM routers choose among models using static per-model costs. We show that open-weight inference markets introduce a second, largely ignored decision axis: after choosing a model, a client must still choose which provider serves it. Measuring live endpoints across [nummodels] open models, competing providers, multiple task types, and three measurement waves, we find that provider choice cannot be inferred from the price list. The same model can vary sharply in quality, latency, availability, and price across providers; higher-priced providers are consistently faster, but price does not reliably predict quality or availability; and provider feasibility is task-selective, with one deployment nearly normal on knowledge tasks but catastrophically degraded on multi-step reasoning. We formulate same-model provider selection as a price-taker market-aware routing problem. A simple measured-map policy routes to the cheapest provider that is both quality-equivalent and healthy, yielding matched-quality savings while avoiding degraded endpoints. Because the map drifts, we introduce FACET, an online provider router that certifies per-(provider x task) feasibility facets and fails safe to an anchor before serving uncertified arms. Across relaxed deployment assumptions, FACET tolerates imperfect task assignment and sparse feedback, while systematic evaluator bias exposes a quality-signal trust boundary that can be mitigated with ground-truth probes or audits. Live provider runs further confirm that certification can move real traffic from a premium anchor to a substantially cheaper certified endpoint. Our results suggest that market-aware LLM routing must measure not only which model to use, but also who serves it.