arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推理拍卖

Inference Auctions

Keegan Harris, Siddharth Prasad, Asher Trockman, Nika Haghtalab, Michael I. Jordan

arXiv 2609.40070首次发表:更新:

发表机构

University of California, Berkeley; Toyota Technological Institute at Chicago; Google Research; Inria & École Normale Supérieure(加州大学伯克利分校; 芝加哥丰田技术学院; 谷歌研究院; 法国国家信息与自动化研究所 & 巴黎高等师范学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对推理需求超过计算能力时的服务优先级问题,提出推理拍卖机制,允许用户竞价获取更快服务,实现经济高效分配且不牺牲延迟,并开发快速定价算法和自动出价代理,实验证明在保持SGLang优势的同时提升系统福利。

AI 中文摘要

当推理需求超过可用计算能力时,模型提供商必须决定哪些请求应优先得到服务。用户对LLM API的延迟有不同的容忍度,但当前的优先定价方案将这些差异压缩为粗略的固定价格服务层级。我们设计了一种推理拍卖,允许用户为更快的服务出价。我们的拍卖以经济高效的方式分配优先级,而不牺牲延迟,并且我们开发了快速算法来实现激励真实出价的价格。我们还为我们的推理拍卖设计了一个自动出价代理,用户指定推理预算,自动出价者随时间动态调整其出价,以在预算约束下最大化用户效用。实验验证了我们拍卖的实用性:它在保持SGLang(一种最先进的推理服务框架)的缓存利用率和延迟优势的同时,提高了系统福利。

英文摘要

When inference demand exceeds available compute capacity, model providers must decide which requests should be served first. Users have different tolerances for delay from an LLM API, but current priority pricing schemes compress these differences into coarse fixed-price service tiers. We design an inference auction that allows users to bid for faster service. Our auction allocates priority in an economically efficient way without sacrificing latency, and we develop fast algorithms for implementing prices that incentivize truthful bidding. We also design an autobidding agent for our inference auction, where users specify an inference budget and the autobidder dynamically adjusts its bids over time to maximize user utility subject to the budget constraint. Experiments validate the practicality of our auction: it increases system welfare while maintaining the cache utilization and latency advantages of SGLang, a state-of-the-art inference serving framework.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑