集中式串行独裁匹配老虎机中的精确遗憾前沿与外部性调度
Exact Regret Frontiers and Externality Scheduling in Centralized Serial-Dictatorship Bandits
- University of Bern(伯尔尼大学)
- EPFL(洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对集中式串行独裁匹配老虎机,揭示外部性下探索配额的精确可达遗憾前沿,并提出估计-求解-跟踪策略以达所有帕累托最优。\n
AI中文摘要:
在集中式串行独裁匹配老虎机中,探索必须使用完整的匹配,因此学习一个玩家-臂对可能会给其他玩家带来遗憾。我们在已知的公共优先级顺序和单位方差的高斯奖励下研究这种外部性。我们证明,匹配层面的Graves-Lai约束可简化为有限多个成对探索配额,并且在顶选分离的实例中,产生一个多项式大小的边际线性规划。在这些实例中,期望对数遗憾系数的精确可达集为$G(\theta)\Xset(\theta)$,其中$\Xset$是可行的匹配分配集,$G$将分配映射到玩家遗憾。通常的上闭Graves-Lai区域尽管具有相同的帕累托最小边界,但可能严格更大。我们进一步表明,相同的探索配额可以通过其调度引发非常不同的遗憾。最后,我们构建估计-求解-跟踪策略,在完整行严格类上一致良好,无需假设优化器唯一性即可达到每个固定的正加权最优。每个帕累托最小点都是逐点可达的,可能通过实例校准的目标实现。
英文摘要:
Exploration in centralized serial-dictatorship matching bandits must use complete matchings, so learning one player-arm pair can impose regret on others. We study this externality under a known common priority order and Gaussian rewards with unit variance. We show that the matching-level Graves-Lai constraints reduce to finitely many pairwise exploration quotas and, at top-choice-separated instances, yield a polynomial-size marginal linear program. At these instances, the exact attainable set of expected logarithmic regret coefficients is $G(θ)\mathcal{X}(θ)$, where $\mathcal{X}$ is the feasible matching-allocation set and $G$ maps allocations to player regret. The usual upper-closed Graves-Lai region can be strictly larger despite having the same Pareto-minimal boundary. We further show that identical exploration quotas can induce very different regret through their scheduling. Finally, we construct estimate-solve-track policies, uniformly good on the full row-strict class, that attain every fixed positively weighted optimum without assuming optimizer uniqueness. Every Pareto-minimal point is pointwise attainable, possibly through an instance-calibrated target.