拟阵约束下的博弈子模极大化:学习压缩交换策略
Bandit Submodular Maximization under Matroid Constraints: Learning Compressed Exchange Policy
浏览论文内容
中文总结 AI 辅助
该研究针对拟阵约束下的对抗博弈子模极大化问题,提出基于平衡分数交换的多项式时间算法,实现次线性regret,是该方向首个此类算法。
中文摘要 AI 辅助
我们研究拟阵约束下单调子模函数的对抗博弈极大化问题。对于含n个元素的秩为k的拟阵,我们提出一种随机预言机多项式时间算法,每轮进行一次可行值查询,其期望(1-1/e) regret为Õ(n^(1/3)k^(2/3)T^(2/3))。这是首个针对一般拟阵约束下对抗博弈子模极大化的次线性regret算法。技术上,我们将该问题视为学习Poisson基游走的交换策略,这将问题与上下文博弈关联并给出信息论次线性regret保证,但直接学习指数级多的策略需要指数时间和空间。因此我们引入“平衡分数交换”,将策略混合压缩为单个分数基,同时保留Poisson分析所需的交换信息,这得到了具有相同regret保证的多项式时间算法。
英文摘要
We study adversarial bandit maximization of monotone submodular functions under a matroid constraint. For a rank-$k$ matroid on $n$ elements, we give a randomized oracle-polynomial algorithm that makes one feasible value query per round and has expected $(1-1/e)$-regret $\widetilde O(n^{1/3}k^{2/3}T^{2/3})$. This is the first sublinear-regret algorithm for adversarial bandit submodular maximization under general matroid constraints. Technically, we view the problem as learning an exchange policy for the Poisson base walk. This connects the problem to contextual bandits and gives an information-theoretic sublinear-regret guarantee, but directly learning the exponentially many policies requires exponential time and space. We therefore introduce \emph{balanced fractional exchanges}, which compress the policy mixture into a single fractional base while retaining the exchange information needed by the Poisson analysis. This leads to an polynomial time algorithm with the same regret guarantee.
发表机构
- Great Bay University(湾大大学)
- Fuzhou University(福州大学)
机构由 AI 辅助整理,请以论文原文为准。