保留、定制还是退出:大语言模型推理服务中的默认设计与代币定价
Keep, Customize, or Exit: Default Design and Token Pricing in LLM Reasoning Services
查看机构详情
- Bilkent University(比尔肯特大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对LLM推理服务的供需交互建立Stackelberg博弈模型,推导最优解,明确默认设置的影响条件,实验验证模型并展示均衡相关因素。
中文摘要 AI 辅助
我们研究一种大语言模型(LLM)服务,其中服务提供商选择每代币价格和默认推理代币分配,用户可选择接受默认设置、定制分配或退出。更大的分配可提升准确率,但会增加代币成本与延迟。我们将此交互建模为Stackelberg博弈,推导出用户唯一最优定制分配的闭式解。对于任意价格,可接受的默认设置要么为空集,要么为紧区间。我们通过三区间规则刻画提供商的最优默认设置,将均衡计算简化为一维价格优化,并证明均衡的存在性。进一步表明,仅当用户重视避免定制的便利性时,默认设置才会影响实际推理分配;否则,所有服务提供结果均会实现用户的最优定制分配。在五个数学与科学基准上,对两个小型开放权重推理模型的实验验证了准确率-代币模型,且展示了模型与任务特征如何决定均衡价格、默认设置及推理分配。
英文摘要
We study a large language model (LLM) service in which a provider chooses a per-token price and a default reasoning-token allocation, while a user may accept the default, customize the allocation, or exit. Larger allocations can improve accuracy but increase token cost and latency. We model this interaction as a Stackelberg game and derive the user's unique optimal customized allocation in closed form. For any price, the acceptable defaults form either an empty set or a compact interval. We characterize the provider's optimal default through a three-regime rule, reduce equilibrium computation to a one-dimensional price optimization, and prove the existence of the equilibrium. We further show that defaults affect the implemented reasoning allocation only when users value the convenience of avoiding customization; otherwise, every service-providing outcome implements the user's optimal customized allocation. Experiments with two compact open-weight reasoning models on five mathematics and science benchmarks support the accuracy-token model and show how model and task characteristics determine equilibrium prices, defaults, and reasoning allocations.