arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09577cs.LG

内生马尔可夫状态下的在线资源分配:更少的LP求解获得更多收益

Online Resource Allocation with an Endogenous Markov State: Fewer LP Solves Earn More

  • Taobao & Tmall Group of Alibaba(阿里巴巴淘宝天猫集团)
  • Renmin University of China(中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhaohua Chen

AI总结:

研究带内生马尔可夫状态的在线资源分配,提出低频重新求解策略,在非退化与退化条件下分别达到O(1)与O(√T)遗憾,并证明频繁优化可能更差。

AI中文摘要:

我们研究有限时域在线资源分配问题,其中请求独立同分布,且存在一个在有限状态空间上的内生马尔可夫状态:每个动作都会影响决定未来奖励和资源消耗的状态转移。在该问题中,一个瞬态流体LP基准为任何非预期策略的期望奖励提供了上界,而一个平稳LP则提供随机化的状态依赖控制。我们假设平稳LP具有唯一最优解,并将原始非退化性和最优诱导核的不可约性识别为该框架中的重要正则条件。在已知请求先验的情况下,我们证明在非退化性和不可约性条件下,频繁和低频重新求解都能达到O(1)的遗憾。然而,在退化最优解下,不可约性使得低频重新求解获得最坏情况下的Θ(√T)速率,而频繁重新求解可能产生Ω(T)的遗憾。因此,更频繁的优化在渐近意义上可能表现更差。在未知请求先验的情况下,我们开发了一种三阶段U形低频重新求解策略,该策略以O(log log T)次LP求解协调学习与库存修正。当最优诱导核不可约且算法被给予最优目标状态类别和一个常数成本进入策略时,它在非退化条件下达到O(1)遗憾,在退化条件下达到O(√T)遗憾。若缺乏目标类别信息,线性极小极大遗憾不可避免。数值实验进一步说明了逐轮重新求解相对于按轮次低频重新求解的不稳定性,表明阈值化能大幅减轻其损失,并发现低频方案在已知和估计先验下均保持优势。

英文摘要:

We study finite-horizon online resource allocation with i.i.d. requests and an endogenous Markov state on a finite state space: each action affects the transition of the state that governs future rewards and resource consumption. In this problem, a transient fluid LP benchmark upper bounds the expected reward of every nonanticipating policy, while a stationary LP supplies randomized state-dependent controls. We assume that the stationary LP has a unique optimum and identify primal nondegeneracy and irreducibility of the optimal induced kernel as important regularity conditions in this framework. With a known request prior, we show that, under nondegeneracy and irreducibility, both frequent and infrequent re-solving attain $O(1)$ regret. However, under a degenerate optimum, irreducibility yields the sharp worst-case $Θ(\sqrt{T})$ rate for infrequent re-solving, while frequent re-solving can incur $Ω(T)$ regret. Thus, more frequent optimization can perform asymptotically worse. With an unknown request prior, we develop a three-phase U-shaped infrequent re-solving policy that coordinates learning and inventory correction with $O(\log\log T)$ LP solves. When the optimal induced kernel is irreducible and the algorithm is given the optimal target state class and a constant-cost entrance policy, it attains $O(1)$ regret under nondegeneracy and $O(\sqrt{T})$ regret under degeneracy. Without the target-class information, linear minimax regret is unavoidable. Numerical experiments further illustrate the instability of round-by-round re-solving relative to epoch-wise infrequent re-solving, show that thresholding greatly mitigates its loss, and find that infrequent schemes remain dominant under both known and estimated priors.

↑