arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12719cs.GTcs.AI

面向大语言模型路由的误差感知反向拍卖机制

Error-Aware Reverse Auction Mechanism for Large Language Model Routing

Haolong Chen, Zhengyuan Xin, Liang Zhang, Lei Xue, Guangxu Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

针对大语言模型路由的信息风险不匹配与可扩展性瓶颈,提出EA-RAM反向拍卖机制,证明其相关特性,实验显示该机制鲁棒且性能优于集中式基线。

中文摘要 AI 辅助

将每个查询路由到具有成本效益的大语言模型(LLM)对于平衡质量与成本至关重要,但大多数路由依赖集中式任务中心预测模型性能,随着模型池扩大,会产生信息风险不匹配和可扩展性瓶颈。我们提出一种基于市场的路由范式,通过反向拍卖将事前预测转移给LLM提供商,提供商以自预测的成功概率和执行成本进行竞价。为考虑提供商预测和中心评估中固有的噪声,我们引入误差感知反向拍卖机制(EA-RAM),该机制明确对这种固有双重误差进行建模。我们证明EA-RAM在双重误差下具有贝叶斯激励兼容性和个体理性,建立了中心理性的充分条件,并推导了明确的福利损失边界。我们进一步确定了鲁棒性效应:符号相反的误差可以抵消,消失尾部链接函数(如逻辑函数)通过饱和稳定明确案例,额外噪声平滑信念图,减少边际操纵的收益。在模拟和真实基准上的实验表明,EA-RAM对双重误差具有鲁棒性,且比集中式基线实现了更好的成本-性能帕累托前沿,当提供商贡献本地信息时还会获得额外收益,验证了其实际有效性。

英文摘要

Routing each query to a cost-effective large language model (LLM) is critical for balancing quality and cost, yet most routers rely on a centralized task center to predict model performance, creating an information-risk mismatch and a scalability bottleneck as the model pool grows. We formulate LLM routing as a market-based allocation problem among strategic providers and propose a routing paradigm that shifts ex-ante prediction to LLM providers via a reverse auction, where providers submit self-predicted acceptance probabilities and execution costs. To account for noisy provider predictions and center evaluations, we introduce the \textit{\textbf{E}rror-\textbf{A}ware \textbf{R}everse \textbf{A}uction \textbf{M}echanism} (EA-RAM), which explicitly models this Dual Error. We prove that, under a private-evaluation-belief structure, truthful effective-surplus reporting is incentive compatible in the reduced-form score space and individually rational under sellers' subjective beliefs, establish sufficient conditions for center rationality, and derive an explicit social-welfare loss bound. We further identify robustness effects: opposite-signed errors can cancel, vanishing-tail link functions (e.g., logistic) stabilize clear-cut cases via saturation, and extra noise smooths belief maps and reduces their maximal local sensitivity. Simulations and real-world benchmarks show that EA-RAM is robust to Dual Error and achieves a better cost--performance Pareto frontier than centralized baselines, with additional gains from provider-side local information, validating its practical effectiveness.

发表机构

  • Shenzhen International Center for Industrial and Applied Mathematics(深圳国际工业与应用数学中心)
  • Shenzhen Research Institute of Big Data(深圳大数据研究院)
  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • Shenzhen Campus of Sun Yat-sen University(中山大学深圳校区)
  • Shenzhen Loop Area Institute(深圳河套学院)

机构由 AI 辅助整理,请以论文原文为准。

↑