arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2510.01867cs.LG

具有约束条件的在线凸优化中的通用动态遗憾和约束违反界

Universal Dynamic Regret and Constraint Violation Bounds for Constrained Online Convex Optimization

  • Indian Institute of Technology, Bombay, India(印度理工学院,孟买,印度)
  • Tata Institute of Fundamental Research, Mumbai, India(塔塔基础研究所在孟买,印度)

机构由 AI 辅助整理,请以论文原文为准。

Subhamon Supantha, Abhishek Sinha

更新

AI总结:

本文提出两种算法,分别解决在线凸优化中的动态遗憾和约束违反问题,改进了现有最先进结果。

AI中文摘要:

我们考虑了在线凸优化(OCO)框架的一个广义扩展,该框架具有对抗性的在线约束。在该问题中,一个在线学习者在多个回合中与对手依次交互。每个回合开始时,学习者从一个凸决策集中选择一个动作。之后,对手揭示一个凸成本函数和一个凸约束函数。学习者的目标是最小化累积成本,同时尽可能紧密地满足约束条件。我们提出了两个高效的算法,具有简单的模块化结构,能够提供通用的动态遗憾界和累积约束违反界,优于现有最先进结果。虽然第一个算法,其达到最优遗憾界,涉及对约束集的投影,第二个算法是无投影的,并在快速变化的环境中实现了更好的违反界。我们的结果在最一般的情况下成立,当成本和约束函数任意选择时,且约束函数不需要包含任何固定的共同可行点。我们通过引入一个通用框架,将受约束的学习问题转化为标准OCO问题的一个实例,其中使用特别构造的替代成本函数来建立这些结果。

英文摘要:

We consider a generalization of the celebrated Online Convex Optimization (OCO) framework with adversarial online constraints. In this problem, an online learner interacts with an adversary sequentially over multiple rounds. At the beginning of each round, the learner chooses an action from a convex decision set. After that, the adversary reveals a convex cost function and a convex constraint function. The goal of the learner is to minimize the cumulative cost while satisfying the constraints as tightly as possible. We present two efficient algorithms with simple modular structures that give universal dynamic regret and cumulative constraint violation bounds, improving upon state-of-the-art results. While the first algorithm, which achieves the optimal regret bound, involves projection onto the constraint sets, the second algorithm is projection-free and achieves better violation bounds in rapidly varying environments. Our results hold in the most general case when both the cost and constraint functions are chosen arbitrarily, and the constraint functions need not contain any fixed common feasible point. We establish these results by introducing a general framework that reduces the constrained learning problem to an instance of the standard OCO problem with specially constructed surrogate cost functions.

补充信息

↑