arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33760cs.LGstat.ML

随机输入与对抗输出下的预言机高效在线分类

Oracle-Efficient Online Classification with Stochastic Inputs and Adversarial Outputs

Gon Buzaglo, Elad Hazan

首次发表
浏览论文内容

中文总结 AI 辅助

针对上下文二元预测,提出带高斯扰动的跟随扰动领导者算法,实现最优遗憾且每轮仅需一次预言机调用,解决开放问题,证明混合分类计算上易处理。

中文摘要 AI 辅助

我们考虑上下文二元预测问题,其中上下文独立同分布地来自未知分布,而损失是自适应选择的。我们证明,对于每个观测到的上下文,使用高斯扰动的简单跟随扰动领导者算法,在包含N个专家的一类问题上,能够实现最优的$\tildetilde O(\tildetilde sqrt{T\tildetilde log N})$期望遗憾,同时每轮仅需一次优化预言机调用,且无需显式枚举该类别。对于无限假设类$\tildetilde mathcal H$,该算法达到$\tildetilde O(\tildetilde sqrt{T\tildetilde operatorname{VC}(\tildetilde mathcal H)})$遗憾。这解决了Lazaric和Munos(2012)提出的一个开放问题,表明混合分类在计算上与统计学习一样容易。

英文摘要

We consider binary prediction with i.i.d. contexts from an unknown distribution and adaptively chosen losses. We show that a simple Follow-the-Perturbed-Leader algorithm using a Gaussian perturbation for each observed context achieves $\widetilde O(\sqrt{T\log N})$ regret for a class of $N$ experts, while requiring one optimization-oracle call per round and no explicit enumeration of the class. For an infinite hypothesis class $\mathcal H$, the same algorithm achieves $\widetilde O(\sqrt{T\operatorname{VC}(\mathcal H)})$ regret. This resolves an open problem posed by Lazaric and Munos (2012), showing that hybrid classification is computationally as easy as statistical learning. As an application, we reduce the problem of contextual bandits with $K$ actions to classification through uniform exploration, achieving $\widetilde O(K^{2/3}T^{2/3}(\log N)^{1/3})$ regret. This matches the best known dependence on the horizon while removing the context-distribution access required by prior oracle-efficient methods.

发表机构

  • Princeton University(普林斯顿大学)
  • Google DeepMind(谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

↑