arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00710cs.DScs.LG

面向具有随机令牌消耗的大语言模型(LLM)API的预测辅助定价与准入

Prediction-Assisted Pricing and Admission for LLM APIs with Stochastic Token Consumption

Patrick Wong

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对LLM API的随机令牌消耗问题,提出PCUCB算法辅助定价与准入,可在全信息与从头学习间平滑插值,保障平台收益与可行性。

中文摘要 AI 辅助

大语言模型(LLM)应用通常会销售或内部分配多种服务产品:小型或高级模型、短或长令牌上限,以及可能的多个挂牌价格。运营决策不仅是选择哪个模型来回答提示,价格会影响购买概率,令牌上限会同时改变用户价值和资源消耗的尾部,而被接受的请求会竞争共享计算资源和高级模型的容量。需求和输出长度最初是不确定的,而离线模型可能提供有用但不完善的预测。我们针对具有随机资源消耗的序列定价与准入问题进行建模:每个到达的请求属于一个可观测的细分段,平台选择一个产品-价格对或不提供报价;购买、收入和资源使用随后是随机的,离线预测器为每个细分段-产品单元提供统一的、经过验证的误差半径。我们提出了预测截断上置信界(Prediction-Clipped UCB,PCUCB)算法,该算法将离线预测区间与在线置信区间相交,使用资源影子价格评估产品,并在承诺前保留样本路径包络;当预测准确时,先验可实现快速启动,而当预测粗略时,在线学习可保护平台。分析是模块化的,在同时置信事件下,针对缓冲流体基准的遗憾由 pacing 项加上相交区间的累积直径界定;对于 J 个细分段-产品单元和预测半径 ε,这给出了遗憾上界为 Õ(√T + (1+Λ̄)min{Tε,√(JT)}),其中 Λ̄ 界定运营影子价格,因此该算法在几乎全信息 regime 和从头学习之间平滑插值,且通过保留包络在每条样本路径上都满足硬可行性。

英文摘要

An LLM application often sells or internally allocates several service products: a small or premium model, a short or long token cap, and possibly multiple posted prices. The operational decision is not merely which model answers a prompt. A price changes purchase probability, a token cap changes both user value and the tail of resource consumption, and accepted requests compete for shared compute and premium-model capacity. Demand and output length are initially uncertain, while an offline model may provide useful but imperfect predictions. We formulate sequential pricing and admission with stochastic resource consumption. Each arriving request belongs to an observable segment. The platform chooses a product--price pair or makes no offer; purchase, revenue, and resource use are then random. An offline predictor supplies a uniform, validated error radius for every segment--product cell. We propose Prediction-Clipped UCB (PCUCB), which intersects the offline prediction interval with an online confidence interval, evaluates products using resource shadow prices, and reserves a sample-path envelope before commitment. The prior gives a fast start when accurate, while online learning protects the platform when predictions are coarse. The analysis is modular. On a simultaneous confidence event, regret against a buffered fluid benchmark is bounded by a pacing term plus the cumulative diameter of the intersected intervals. For $J$ segment-product cells and prediction radius $\varepsilon$, this yields \[ \widetilde O\left( \sqrt{T}+(1+\barΛ) \min\{T\varepsilon,\sqrt{JT}\} \right), \] where $\barΛ$ bounds operational shadow prices. Thus the algorithm smoothly interpolates between an almost full-information regime and learning from scratch. Hard feasibility holds on every sample path through reservation envelopes.

↑