arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00662cs.AI

面向稀疏上下文与共享预算的感知漂移大语言模型路由

Drift-Aware LLM Routing with Sparse Contexts and Shared Budgets

Cheung Hao Lee, Patrick Wong

首次发表
浏览论文内容

中文总结 AI 辅助

针对多模型语言服务路由的预算约束与漂移问题,提出DRS算法,通过滚动审计窗口估计并结合悲观奖励等策略实现路由,理论分析了其遗憾界与自适应速率。

中文摘要 AI 辅助

多模型语言服务需在路由每个请求的同时,维持计算、延迟、内存或资金成本等工作负载级预算。两个特征使该问题比静态模型选择困难得多:提示表示维度极高,仅一小部分嵌入方向可预测模型的增量价值;且请求组合、模型前沿会在发布、微调、量化变更及系统更新后发生漂移。我们将该问题形式化为带有多重背包约束的非平稳稀疏上下文路由,以及可选的影子审计流——该流会在多个模型上评估一小部分提示。我们提出感知漂移稀疏路由(DRS),其策略通过滚动审计窗口估计奖励与资源使用,采用悲观奖励和乐观成本估计进行路由,在线更新资源影子价格,并在提交前应用硬计量器。分析将控制与统计分离,在具有均匀预测半径{β_t}的任意事件中,相对于 paced 动态流体基准的遗憾被约束为半径之和、容量缓冲项及O(√T) pacing项的总和。在稀疏线性模型和有界漂移V_T下,滚动估计给出\\( \widetilde O\left( T\sqrt{\frac{s}{\rho W}}+WV_T+\sqrt{T} \right) \\),其中s为稀疏度,ρ为审计率,W为窗口长度。优化W后,当V_T=0时得到常规平稳O(√(sT/ρ))速率,漂移下得到O(T^{2/3}(s/ρ)^{1/3}V_T^{1/3})的自适应项。

英文摘要

A multi-model language service must route each request while preserving workload-level budgets for compute, latency, memory, or monetary cost. Two features make this problem materially harder than static model selection. Prompt representations are high dimensional, so only a small subset of embedding directions may predict the incremental value of a model, and both the request mix and the model frontier drift after launches, fine-tunes, quantization changes, and system updates. We formulate nonstationary sparse contextual routing with multiple knapsack constraints and an optional shadow-audit stream that evaluates a small fraction of prompts on several models. We propose Drift-Aware Sparse Routing (DRS). The policy estimates reward and resource use from a rolling audit window, routes using pessimistic reward and optimistic cost estimates, updates resource shadow prices online, and applies a hard meter before commitment. The analysis separates control from statistics. On any event with uniform prediction radii $\{β_t\}$, regret against a paced dynamic fluid benchmark is bounded by the sum of the radii, a capacity-buffer term, and an $O(\sqrt{T})$ pacing term. Under a sparse linear model and bounded drift $V_T$, rolling estimation gives \[ \widetilde O\left( T\sqrt{\frac{s}{ρW}}+WV_T+\sqrt{T} \right), \] where $s$ is sparsity, $ρ$ is the audit rate, and $W$ is the window length. Optimizing $W$ yields the usual stationary $O(\sqrt{sT/ρ})$ rate when $V_T=0$ and a $O(T^{2/3}(s/ρ)^{1/3}V_T^{1/3})$ adaptation term under drift.

↑