arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20968cs.LG

从切换遗憾到动态遗憾:一种通过无偏随机序列的简单归约

From Switching to Dynamic Regret: A Simple Reduction via Unbiased Random Sequences

Yibo Wang, Wenhao Yang, Sifan Yang, Wei Jiang, Yuanyu Wan, Lijun Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出一种简单归约框架,将动态遗憾最小化转化为切换遗憾最小化,通过构造无偏随机序列,为强凸、指数凹和一般凸损失分别建立了匹配极小极大最优的动态遗憾界。

中文摘要 AI 辅助

在非平稳在线学习中,动态遗憾作为衡量在线学习者在面对随时间变化的比较器序列时表现优劣的指标,日益受到关注。尽管取得了相当大的进展,但对于强凸和指数凹损失函数,获得最优界通常涉及复杂的分析。在本文中,我们提出了一个简单的框架,将动态遗憾最小化归约为切换遗憾最小化。因此,我们可以通过使用具有切换遗憾保证的现成算法来推导动态遗憾界。我们归约的关键思想是,对于任何比较器序列,构造一个辅助随机序列,该序列在每一轮都是无偏的,具有受控的方差和可控的切换次数。将此构造与合适的替代损失相结合,我们可以将动态遗憾分解为针对随机序列的期望切换遗憾及其受控方差。理论上,对于强凸和指数凹损失,我们建立了$\widetilde{O}(T^{1/3}P_T^{2/3})$的动态遗憾界,其中$T$表示时间范围,$P_T$表示比较器序列的路径长度。此外,对于一般凸损失,相同的归约也恢复了$O(\sqrt{T(1+P_T)})$的动态遗憾界。值得注意的是,我们所有的结果都匹配这三种损失类型的极小极大最优结果,凸显了我们所提出框架的多功能性。

英文摘要

In non-stationary online learning, dynamic regret has attracted increasing attention as a measure of how well an online learner performs against a time-varying comparator sequence. Despite considerable advances, attaining optimal bounds for strongly convex and exp-concave losses often involves intricate analysis. In this paper, we present a \textit{simple} framework that reduces dynamic regret minimization to switching regret minimization. As a result, we can derive dynamic regret bounds by using off-the-shelf algorithms with switching regret guarantees. The key idea of our reduction is to construct, for \textit{any} comparator sequence, an auxiliary random sequence that is unbiased at each round, with the controlled variance and a manageable number of switches. Combining this construction with suitable surrogate losses, we can decompose dynamic regret into the expected switching regret against the random sequence and its controlled variance. Theoretically, for strongly convex and exp-concave losses, we establish the $\widetilde{O}(T^{1/3}P_T^{2/3})$ dynamic regret bounds, where $T$ denotes the time horizon and $P_T$ denotes the path-length of the comparator sequence. Moreover, for general convex losses, the same reduction also recovers the $O(\sqrt{T(1+P_T)})$ dynamic regret bound. Notably, all our findings match the minimax optimal results for these three types of losses, highlighting the versatility of our proposed framework.

发表机构

  • State Key Laboratory for Novel Software Technology, Nanjing University(南京大学计算机软件新技术全国重点实验室)
  • School of Artificial Intelligence, Nanjing University(南京大学人工智能学院)
  • School of Software Technology, Zhejiang University(浙江大学软件学院)

机构由 AI 辅助整理,请以论文原文为准。

↑