一种具有常数遗憾的在线资源分配一阶学习算法
A First-Order Learning Algorithm for Online Resource Allocation with Constant Regret
浏览论文内容
中文总结 AI 辅助
针对有限时域在线资源分配问题,提出一种原始一阶学习策略,每周期仅做一次梯度上升更新,实现常数遗憾且无需线性规划求解和非退化假设。
中文摘要 AI 辅助
我们研究了一个有限时域在线资源分配问题,其中初始资源容量与时间范围成比例。在每个周期内,观察到一种请求类型,并从有限菜单中选择一个动作。每个动作获得奖励并消耗一组资源向量。到达类型是独立同分布的,但其概率未知。我们提出了一种原始一阶学习策略,该策略在每个周期内对与当前请求类型相关的动作坐标执行一次梯度上升更新。该策略相对于事后最优实现了$O(1)$的期望加性遗憾,且其界与时间范围$T$无关。它不需要求解任何线性规划,并且遗憾界不要求流体线性规划的非退化假设。
英文摘要
We study a finite-horizon online resource allocation problem with initial resource capacities proportional to the horizon. In each period, a request type is observed and one action is chosen from a finite menu. Each action earns a reward and consumes a vector of resources. The arrival types are independent and identically distributed, but their probabilities are unknown. We present a primal first-order learning policy that, in each period, performs one gradient ascent update of the action coordinates associated with the current request type. The policy achieves $O(1)$ expected additive regret relative to the hindsight optimum, with a bound independent of the horizon $T$. It does not solve any linear program, and the regret bound does not require a nondegeneracy assumption on the fluid linear program.
发表机构
- City University of Hong Kong(香港城市大学)
- New York University(纽约大学)
机构由 AI 辅助整理,请以论文原文为准。