arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具有静态和时变曝光下限的差异舍入公平博弈

Discrepancy-Rounded Fair Bandits with Static and Time-Varying Exposure Floors

Ibne Farabi Shihab, Joyanta Jyoti Mondal, Anuj Sharma

arXiv 2607.22935首次发表:更新:

AI 中文总结

研究具有曝光下限的随机博弈舍入问题,提出分块模型,如 BDQ-UCB 等算法,在不同场景下有相应遗憾界,实验验证算法在合成数据等测试中能精确可行且遗憾与基线有竞争力。

AI 中文摘要

在推荐、内容策划和规范分配中,当每个提供者、策略或组必须在一段时间内获得保证曝光而不仅仅是总体曝光时,会出现最小曝光约束。我们研究了具有精确曝光下限的随机博弈,并表明正确的对象是一个舍入问题:分数公平调度通过整数拉动来实现,曝光误差恰好是一个差异向量。主要贡献是一个具有时变下限的分块模型。BDQ-UCB 确定性地满足每个块下限,并且具有由非强制性预算 $R$ 而非时间范围 $T$ 控制的公平遗憾,高概率遗憾为 $O(\sqrt{KR\log(KT)})$。MOSS 残差变体达到 $O(\sqrt{KR})$,匹配的下限给出了极小极大率 $\Theta(\sqrt{KR})$,即使有正的强制曝光;kl-UCB$^{++}$ 残差规则增加了实例相关的最优性。对于重叠组下限,该公式变得至关重要:按策略舍入可能会在组大小上违反组约束 $\Omega(s)$,而 Beck--Fiala 零空间舍入在块预算内满足每个组下限,违反低于策略度 $t$,并与具有相同 $R$ 参数化遗憾的 UCB 组合。对于学习到的组计划,我们在 $\widetilde\Theta(\sqrt{KT})$ 处关闭不相交系统,给出一个双账本分解来解释为什么朴素索引规则在重叠下失败,并证明一个在初始覆盖松弛条件下路径可行且达到条件 $\widetilde O(\sqrt{KT})$ 保证的计划采样规则,无条件重叠率未解决。在合成下限、MovieLens-100k 类型曝光和部署压力测试上的实验表明,无需惩罚调整即可实现精确可行性,并且遗憾与调整后的拉格朗日基线具有竞争力。

英文摘要

Minimum-exposure constraints arise in recommendation, content curation, and regulated allocation when each provider, arm, or group must receive guaranteed exposure inside a period rather than only in aggregate. We study stochastic bandits with exact exposure floors and show that the right object is a rounding problem: a fractional fair schedule is realized as integral pulls, and the exposure error is exactly a discrepancy vector. The main contribution is a blockwise model with time-varying floors. BDQ-UCB satisfies every block floor deterministically and has fair regret governed by the nonmandatory budget $R$, not the horizon $T$, with high-probability regret $O(\sqrt{KR\log(KT)})$. A MOSS residual variant attains $O(\sqrt{KR})$, and a matching lower bound gives the minimax rate $Θ(\sqrt{KR})$, even with positive mandatory exposure; a kl-UCB$^{++}$ residual rule adds instance-dependent optimality. The formulation becomes essential for overlapping group floors: per-arm rounding can violate a group constraint by $Ω(s)$ in the group size, whereas Beck--Fiala null-space rounding meets every group floor within the block budget with violation below the arm degree $t$, and composes with UCB at the same $R$-parametrized regret. For learned group plans, we close disjoint systems at $\widetildeΘ(\sqrt{KT})$, give a dual-ledger decomposition explaining why naive index rules fail under overlap, and prove a plan-sampling rule that is pathwise feasible under an initial cover-slack condition and attains a conditional $\widetilde O(\sqrt{KT})$ guarantee, leaving the condition-free overlap rate open. Experiments on synthetic floors, MovieLens-100k genre exposure, and deployment stress tests show exact feasibility without penalty tuning and regret competitive with tuned Lagrangian baselines.

Comments28 pages, 8 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑