发表机构
Princeton University; Google Research(普林斯顿大学; 谷歌研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对对抗环境下预算约束在线学习问题,提出全信息算法,匹配已有下界,可扩展至在线资源分配,是首个能为相关任务实现o(√T)保证的算法。
AI 中文摘要
我们在对抗环境中,针对任意类别的预算节奏专家,为预算约束的在线学习建立了近最优的遗憾界。具体而言,给定任意F个专家的类别和一个候选预算节奏调度表,我们提供了一种全信息算法,该算法对所有累计支出与该调度表距离不超过D的专家,可获得遗憾O(D√log F + √(T log F)),与Braverman等人(2025)建立的下界匹配。我们还证明,该技术可扩展至在线资源分配中的各类问题,在这类问题中,学习者可看到当前可用选项的奖励与成本,且当允许分数分配时,可建立O(D√log F)的遗憾界。据我们所知,这是首个能为此类任务实现o(√T)保证的算法。
英文摘要
We establish near-optimal regret bounds for budget-constrained online learning against arbitrary classes of budget-pacing experts in the adversarial setting. In particular, given any class of $F$ experts and a candidate budget pacing schedule, we provide a full-information algorithm which obtains regret $O(D \sqrt{\log F}+ \sqrt{T\log F})$ against all experts whose cumulative spending stays within distance $D$ of this schedule, matching lower bounds established by Braverman et al. (2025). We additionally show that our technique extends to various problems in online resource allocation, where the learner gets to see the rewards and costs of the current options available to them, and establish $O(D\sqrt{\log F})$ regret bounds when fractional allocation is allowed. This is the first algorithm we are aware of which can achieve $o(\sqrt{T})$ guarantees for such tasks.