为时间定价,而非仅为代币:面向LLM推理的延迟感知机制设计
Pricing Time, Not Just Tokens: Latency-Aware Mechanism Design for LLM Inference
浏览论文内容
中文总结 AI 辅助
针对LLM推理定价,提出延迟感知机制设计,通过分离定理将三维筛选简化为分层一维筛选,并验证两层定价可提升26-66%利润。
中文摘要 AI 辅助
LLM定价的经济理论将代币视为同质商品,将代币总数视为买卖双方考虑的主要特征。我们将推理建模为服务市场,其中买方具有三维私有信息——支付意愿、任务量和时间偏好——且效用取决于延迟余量以及代币数量。我们的主要结果是分离定理:离散硬件层级诱导对时间偏好的内生自选择,将三维筛选简化为每个层级内的标准一维筛选。我们从GPU推理物理中推导出成本结构——计算受限的预填充和带宽受限的解码——并通过虚拟价值技术刻画最优分层机制。最优的每任务价格与数量无关,为平坦的每代币API定价提供了理论基础。我们通过校准至8-GPU H100和B200硬件集群来实证验证该机制。分离定理在105个测试配置中的83%中成立,在经济相关的支付意愿尺度下上升至96%。采用最优机制下两层定价的卖方比最佳单层替代方案多获得26-66%的利润,其收益由硬件成本占每请求价值显著比例的机制中高效的跨层分配驱动。
英文摘要
The economic theory of LLM pricing treats tokens as a homogeneous commodity considering aggregate token count as the main features buyers and sellers consider. We model inference as a service market where buyers have three-dimensional private information - willingness-to-pay, task volume, and time preference - and utility depends on latency slack alongside token quantities. Our main result is a separation theorem: discrete hardware tiers induce endogenous self-selection on time preferences, reducing three-dimensional screening to standard one-dimensional screening within each tier. We derive the cost structure from GPU inference physics - compute-bound prefill and bandwidth-bound decode - and characterize optimal tiered mechanisms via virtual-value techniques. Optimal per-task prices are volume-independent, providing theoretical grounding for flat per-token API pricing. We verify the mechanism empirically by calibrating to 8-GPU clusters of H100 and B200 hardware. The separation theorem holds in 83% of 105 tested configurations overall, rising to 96% at economically relevant WTP scales. A seller adopting two-tier pricing under the optimal mechanism captures 26-66% higher profit than the best single-tier alternative, with gains driven by efficient cross-tier allocation in regimes where hardware costs are a significant fraction of per-request value.
发表机构
- University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
- NVIDIA(英伟达)
机构由 AI 辅助整理,请以论文原文为准。