发表机构
The University of Arizona(亚利桑那大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究具有通用奖励函数逼近的上下文组合半贪心问题,提出计算高效算法,通过解决凸优化问题采样组合动作平衡探索与利用,证明其实现\(O(\sqrt{m A T \log |\mathcal{F}|})\)极小极大最优遗憾界,推广到多种设置。
AI 中文摘要
我们研究了具有通用奖励函数逼近的上下文组合半贪心(CCSB)问题。在每一轮中,学习者观察一个上下文,选择一个由基本臂的子集组成的组合动作,并获得每个所选臂的奖励;目标是随着时间的推移最大化累积奖励。我们提出了一种计算高效的算法,在每一轮中,通过解决一个凸优化问题来采样一个平衡探索和利用的组合动作。该算法可扩展到大型臂集,并且除了对每个组合动作的基数限制为\(m\)之外,对动作集不做任何结构假设。我们证明该算法实现了\(O(\sqrt{m A T \log |\mathcal{F}|})\)的极小极大最优遗憾界,其中\(A\)是臂的数量,\(m\)是组合动作中臂的最大数量,\(T\)是时间范围,\(\mathcal{F}\)是奖励函数类。在可实现的设置中,这个界与基于策略搜索的算法在更受限的候选推荐设置中所达到的最优遗憾保证相匹配,同时推广到任意组合动作结构和通用奖励函数逼近。
英文摘要
We study the contextual combinatorial semi-bandit (CCSB) problem with general reward function approximation. At each round, the learner observes a context, selects a combinatorial action consisting of a subset of basic arms, and receives the reward of each selected arm; the goal is to maximize the cumulative reward over time. We propose SquareCB.Comb, a computationally efficient algorithm that, at each round, solves a convex optimization problem to sample a combinatorial action that balances exploration and exploitation. SquareCB.Comb scales to large arm sets and imposes no structural assumptions on the action set beyond a cardinality bound of $m$ on each combinatorial action. We prove that SquareCB.Comb achieves a minimax optimal regret bound of $O(\sqrt{m A T \log |\mathcal{F}|})$, where $A$ is the number of arms, $m$ is the maximum number of arms in a combinatorial action, $T$ is the time horizon, and $\mathcal{F}$ is the reward function class. In the realizable setting, this bound matches the state-of-the-art regret guarantees achieved by policy search-based algorithms in the more restricted slate recommendation settings, while simultaneously generalizing to arbitrary combinatorial action structures and general reward function approximation.
Comments59 pages (11 pages main body, 17 pages supplementary materials)