arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向多智能体系统的无投影多臂老虎机在线优化:动态遗憾视角

Projection-Free Bandit Online Optimization for Multi-Agent Systems with Dynamic Regret

Xia Jiang, Lu Liu, Gang Feng

arXiv 2608.30159首次发表:更新:

发表机构

Nanyang Technological University; City University of Hong Kong(南洋理工大学; 香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对多智能体动态系统的在线优化问题,提出仅依赖实时数据的分布式无投影多臂老虎机在线优化算法,通过仿真验证其有效性并建立动态遗憾界。

AI 中文摘要

本文研究具有受限输入和时变代价函数的多智能体动态系统的分布式在线优化问题。在线凸优化为序列决策提供了主要框架,但现有在线学习与优化算法通常需要精确的系统模型,限制了其在实际场景中的适用性。为克服这一挑战,本文提出一种分布式多臂老虎机在线反馈优化算法,该算法仅依赖实时输入输出数据。该算法采用平滑零阶单点估计器,直接从代价评估中构建局部梯度近似;此外,为有效实施输入约束,集成了无投影条件梯度更新,使算法适用于在线和大规模场景。本文还建立了依赖于系统非平稳性时间变化度量的次线性动态遗憾界,最后通过数值仿真验证了所提算法的有效性。

英文摘要

This paper investigates distributed online optimization for multi-agent dynamical systems with constrained inputs and time-varying cost functions. While online convex optimization offers a principal framework for sequential decision-making, existing online learning and optimization algorithms typically require accurate system models, limiting their applicability in practical settings. To overcome this challenge, we propose a distributed bandit online feedback optimization algorithm that relies solely on real-time input-output data. The algorithm employs a smoothing zeroth-order one-point estimator to construct local gradient approximations directly from cost evaluations. Additionally, to enforce input constraints effectively, we integrate a projection-free conditional gradient update, making the algorithm well-suited for online and large-scale settings. Furthermore, we establish a sublinear dynamic regret bound that depends on a temporal variation measure of system non-stationarity. Finally, numerical simulations demonstrate the effectiveness of the proposed algorithm.

Comments10 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑