arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在上下文相关服务速率下调度作业时的学习:一种任意时刻速率最优算法

Learning While Scheduling Jobs under Context-Dependent Service Rates: An Anytime Rate-Optimal Algorithm

Seoungbin Bae, Dabeen Lee

arXiv 2610.06006首次发表:更新:

发表机构

KAIST; Seoul National University(韩国科学技术院; 首尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出WISE算法,在未知服务速率和时域下实现上下文排队老虎机的速率最优队列长度遗憾,并给出首个容量松弛相关下界。

AI 中文摘要

我们研究上下文排队老虎机问题,其中学习者在调度作业的同时学习由作业-服务器特征的逻辑函数建模的未知服务速率。性能通过队列长度遗憾来衡量,即相对于知道服务速率的预言机,在第 $t$ 轮时的期望超额队列长度。现有的衰减遗憾保证要么具有次优的衰减速率,要么需要已知的固定时间范围。它们还假设上下文级松弛和特征协方差的严格正最小特征值。在本文中,我们提出WISE(最宽区间选择与消除),在不知道时间范围的情况下,在每一个足够大的时刻实现速率最优的$\tilde{\mathcal{O}}(t^{-1/2})$队列长度遗憾。我们假设容量松弛,即最佳服务器服务下的期望到达工作负载低于服务容量,并且不施加协方差下界。我们的分析使用一个工作负载势,衡量等待作业在其最佳服务器上所需的期望服务尝试次数。其在非空轮次上的漂移结合了由容量松弛确保的负项和次优服务选择产生的误差。然后,一个椭圆势计数限制了WISE选择宽置信区间的频率,从而限制了具有大服务误差的轮次数。我们还改进了现有下界对到达率的依赖,并使其对特征维度和服务器数量的依赖显式化。我们证明了另一个下界,量化了遗憾随归一化容量松弛减小的增加;据我们所知,这是CQB的第一个此类下界。模拟表明,即使上下文级松弛失败,遗憾也很小。

英文摘要

We study contextual queueing bandits, where a learner schedules jobs while learning unknown service rates modeled by logistic functions of job-server features. Performance is measured by queue length regret, the expected excess queue length at round $t$ relative to an oracle that knows the service rates. Existing decaying-regret guarantees either have a suboptimal decay rate or require a known fixed horizon. They also assume context-wise slack and a strictly positive minimum eigenvalue of the feature covariance. In this paper, we propose WISE (Widest Interval Selection with Elimination), achieving rate-optimal $\widetilde{\mathcal O}(t^{-1/2})$ queue length regret at every sufficiently large time without knowing the horizon. We assume capacity slack, meaning that expected incoming workload under best-server service is below service capacity, and impose no covariance lower bound. Our analysis uses a workload potential measuring the expected service attempts needed by waiting jobs on their best servers. Its drift on nonempty rounds combines a negative term ensured by capacity slack with errors from suboptimal service choices. Then an elliptical potential count bounds how often WISE selects wide confidence intervals, thereby limiting the number of rounds with large service errors. We also sharpen the arrival-rate dependence of an existing lower bound and make its dependence on feature dimension and server count explicit. We prove another lower bound that quantifies the increase in regret as the normalized capacity slack decreases; to our knowledge, this is the first such lower bound for CQB. Simulations show small regret even when context-wise slack fails.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑