arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11383cs.LGmath.OC

带背包约束的上下文老虎机的再优化算法

Reoptimization Algorithms for Contextual Bandits with Knapsack Constraints

  • University of Liverpool(利物浦大学)
  • Management School, University of Liverpool(利物浦大学管理学院)

机构由 AI 辅助整理,请以论文原文为准。

Zhen Xu

AI总结:

针对带背包约束的上下文老虎机问题,本文提出UCB算法族的扩展再优化算法,将平均遗憾界从相关动态定价问题的$O(\frac{1}{\text{sqrt}(T)})$降至$O(\frac{(\text{ln } T)^3}{T})$,提升了算法性能。

AI中文摘要:

我们研究带背包约束的上下文老虎机的新算法。在这类问题中,存在有限类型的顾客、产品和资源;每种产品由固定组合的资源制成,且资源具有有限容量。决策者必须为每个到达的顾客从一组多个可能产品中分配一种,将顾客分配到产品会产生随机奖励,该奖励等于顾客与产品特征的未知线性函数加上噪声项。目标是联合学习平均奖励函数,并进行在线分配,以相对于知晓奖励函数的最优策略最小化预期收益损失。我们提出了上置信界(Upper-Confidence-Bound,UCB)算法族的自然且简单的扩展,并应用再优化技术。我们证明,利用再优化,我们的算法可实现平均遗憾为$O(\frac{(\text{ln } T)^3}{T})$,其中$T$是时间范围长度;该界显著降低了文献中基于再优化的密切相关动态定价问题的$O(\frac{1}{\text{sqrt}(T)})$界。

英文摘要:

We study new algorithms for Contextual Bandits with Knapsack. In these problems, there are finitely many types of customers, products, and resources. Each product is made from a fixed combination of resources, and resources have finite capacity. A decision maker must assign each arriving customer one out of a set of multiple possible products. Every assignment of a customer to a product will generate a random reward, which equals an unknown linear function of customer and product features, plus a noise term. The objective is to jointly learn the mean reward function, and to make online assignments to minimize the expected revenue loss relative to an optimal policy that knows the reward function. We propose a natural and simple extension of the Upper-Confidence-Bound (UCB) family of algorithms and apply re-optimization techniques. We show that by taking advantage of re-optimization, our algorithm achieves an average regret of $O(\frac{(\ln T)^3}{T})$ where $T$ is the horizon length. Our bound significantly reduces the $O(\frac{1}{\sqrt{T}})$ bound in the literature for closely related dynamic-pricing problems that are based on re-optimization.

↑