arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Oracle高效且无参数的不可知平滑在线学习

Oracle-Efficient and Parameter-Free Agnostic Smoothed Online Learning

Sasha Voitovych, Adam Block, Alexander Rakhlin, Abhishek Shetty

arXiv 2610.10499首次发表:更新:

发表机构

MIT; Columbia University; Georgia Tech(麻省理工学院; 哥伦比亚大学; 佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出首个无需基测度知识、无参数且Oracle高效的不可知平滑在线学习算法,基于高斯跟随扰动领导者,实现次线性遗憾,填补了理论与实用间的空白。

AI 中文摘要

在线学习在许多领域中是一个有吸引力的框架,因为它即使在数据具有依赖性或是被对抗性选择的情况下,也能实现定义良好的学习。然而,这种通用性带来了高昂的代价,引入了显著的统计和计算障碍。最近,平滑在线学习作为一个有前景的框架出现,它通过假设每个协变量的条件分布相对于某个固定基测度$\mu$具有至多$1/\sigma$的密度,在完全对抗性和完全随机设置之间进行插值,并且已知能够匹配经典学习的统计和计算保证,同时仍然允许在线学习的许多灵活性。然而,现有的Oracle高效算法需要(i)对基测度$\mu$的采样访问,或(ii)由固定假设完美预测的标签。这两个假设都限制了这些算法的适用性,这与统计学习形成对比,在统计学习中,经验风险最小化(ERM)在不可知设置下无需任何数据分布知识即可高效学习。我们证明这两个假设都不是必要的,给出了第一个在不可知设置下无需$\mu$知识即可实现次线性遗憾的Oracle高效算法。我们的算法基于高斯跟随扰动领导者(Gaussian Follow-The-Perturbed-Leader),是无参数的:它不需要知道$\mu$、平滑参数$\sigma$或时间范围$T$,并且对于VC维度为$d$的二元类别,通过每轮单次调用ERM Oracle即可实现$\widetilde O(d\sqrt{T/\sigma})$的遗憾,这在$\sqrt{d}$因子内是最优的。在建立遗憾界的过程中,我们引入了几个可能具有独立兴趣的新技术。

英文摘要

Online learning is an attractive framework in many domains because it permits well-defined learning even when data are dependent or chosen adversarially. This generality, however, comes at a steep price, introducing significant statistical and computational barriers. Recently, smoothed online learning has emerged as a promising framework that interpolates between the fully adversarial and fully stochastic settings by assuming that the conditional law of each covariate has density at most $1/σ$ with respect to some fixed base measure $μ$, and it is known to match the statistical and computational guarantees of classical learning while still allowing for much of the flexibility of online learning. However, existing oracle-efficient algorithms require either (i) sampling access to the base measure $μ$ or (ii) labels that are perfectly predicted by a fixed hypothesis. Both assumptions limit the applicability of these algorithms, in contrast to statistical learning, where empirical risk minimization (ERM) learns efficiently in the agnostic setting without any knowledge of the data distribution. We show that neither assumption is necessary, giving the first oracle-efficient algorithm that achieves sublinear regret in the agnostic setting without knowledge of $μ$. Our algorithm, based on Gaussian Follow-The-Perturbed-Leader, is parameter-free: it requires no knowledge of $μ$, the smoothing parameter $σ$, or the horizon $T$, and it achieves regret $\widetilde O(d\sqrt{T/σ})$ for binary classes of VC dimension $d$ with a single call to an ERM oracle per round, which is optimal up to a $\sqrt{d}$ factor. En route to establishing the regret bound, we introduce several new techniques that may be of independent interest.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑