arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在学习上下文相关服务速率时的非抢占式调度

Nonpreemptive Scheduling While Learning Context-Dependent Service Rates

Wansoo Choi, Seoungbin Bae, Dabeen Lee

arXiv 2609.37660首次发表:更新:

发表机构

Seoul National University; KAIST(首尔大学; 韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对单服务器非抢占式上下文排队老虎机,提出学习-清理-规划(LCP)算法,利用贝尔曼递归实现近最优队列长度遗憾,并在未知时间范围时采用估计SEPT策略跟踪最优。

AI 中文摘要

我们研究单服务器系统中的非抢占式上下文排队老虎机问题。每个作业由一个$d$维上下文向量表示;在每一轮中,一个作业可能到达,其上下文从未知分布$\mathcal{D}$中抽取,其离开概率由该上下文向量的逻辑斯蒂模型决定,其中包含未知参数$\theta^*$。服务器在决定服务哪个等待作业以及是否空闲的同时,从服务结果中学习,旨在最小化队列长度遗憾,即其期望最终队列长度与可行策略可达到的最小值之间的差距。一旦被选择,作业必须被服务至完成,我们称之为非抢占式设置。一个核心挑战是,即使完全了解模型,最优策略通常也不能用简单的短视规则来刻画,因为在相同的队列状态下,最优行动可能随剩余时间范围而变化。然而,当模型和时间范围已知时,最优行动可以通过有限时域贝尔曼递归获得。受此启发,我们提出了学习-清理-规划(LCP)算法,该算法估计系统并使用由此产生的贝尔曼递归来做出依赖于时间范围的决策。LCP实现了$\widetilde{O}(\sqrt{d/T})$的队列长度遗憾,而一个下界构造给出了在某个实例上每个学习策略的$\Omega(\min\{1/\sqrt{d},\sqrt{d/T}\})$遗憾,从而在$T\ge d^2$时建立了至多对数因子意义上的最优性。当时间范围未知时,没有任何独立于时间范围的策略能够针对有限时域最优值实现消失遗憾。因此,我们使用SEPT(即服务具有最高离开概率的等待作业的策略)作为固定参考,并提出一种估计的SEPT算法,该算法在不知道模型的情况下实现了$\widetilde{O}(\sqrt{d/t})$的跟踪误差。

英文摘要

We study nonpreemptive contextual queueing bandits in a single-server system. Each job is represented by a $d$-dimensional context vector; in each round, a job may arrive with its context drawn from an unknown distribution $\mathcal{D}$, and its departure probability is determined by a logistic model of that context vector with an unknown parameter $θ^*$. The server learns from service outcomes while deciding which waiting job to serve and whether to idle, aiming to minimize queue-length regret, the gap between its expected terminal queue length and the minimum achievable by an admissible policy. Once selected, a job must be served until completion, and we refer to this as the nonpreemptive setting. A central challenge is that, even with full model knowledge, the optimal policy cannot in general be characterized by a simple myopic rule, since the optimal action can change with the remaining horizon at the same queue state. Nevertheless, when the model and horizon are known, the optimal action can be obtained through a finite-horizon Bellman recursion. Motivated by this, we propose Learn--Clear--Plan (LCP), which estimates the system and uses the resulting Bellman recursion to make horizon-dependent decisions. LCP achieves $\widetilde{O}(\sqrt{d/T})$ queue-length regret, while a lower-bound construction gives $Ω(\min\{1/\sqrt{d},\sqrt{d/T}\})$ regret for every learning policy on some instance, establishing optimality up to polylogarithmic factors when $T\ge d^2$. When the horizon is unknown, no horizon-independent policy achieves vanishing regret against the finite-horizon optimum. We therefore use SEPT, the policy that serves a waiting job with the highest probability of departure, as a fixed reference, and suggest an estimated-SEPT algorithm that achieves a tracking error of $\widetilde{O}(\sqrt{d/t})$ without knowing the model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑