arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37932cs.LG

学习何时更新:一种近最优的时机老虎机方法

Learning When to Update: A Near-Optimal Timing Bandit Approach

Qiulin Lin, Junyan Su, Liyuan Wang, Minghua Chen

首次发表
浏览论文内容

中文总结 AI 辅助

针对动态系统中未知退化模式的更新时机优化问题,提出时机老虎机框架及BCAE算法,实现近最优遗憾并匹配理论下界。

中文摘要 AI 辅助

在动态环境中运行的系统需要及时更新以维持性能。对于资源密集型系统,如机器学习模型和数字孪生,战略性地安排更新时机至关重要。更新过于频繁会浪费资源,而更新过少则会导致代价高昂的性能下降。当系统的退化模式事先未知时(这在新的运行环境中很常见),该问题尤其具有挑战性。我们将这一挑战形式化为一个新的“时机老虎机”问题,其中每个臂代表一个候选更新间隔,具有固定的更新成本和未知的随机退化成本。三个结构性属性使该设置区别于标准的多臂老虎机:选择一个间隔会使学习者在下次更新前承诺多个时隙;臂成本由每步退化成本和固定更新成本组成;选择更长的间隔自然会在每个中间步骤揭示退化,提供与较短间隔相关的连续反馈。通过利用这些结构,我们开发了平衡连续臂消除(BCAE)算法。BCAE实现了$\tilde{O}(\sqrt{T})$的遗憾,在此设置中优于标准老虎机算法的$\tilde{\Omega}(K\sqrt{T})$遗憾,其中$K$是候选更新间隔的数量。我们进一步提出了一种乐观增强变体(OE-BCAE),该变体整合了置信下界原理以提高经验适应性,同时保持相同的遗憾阶。此外,我们算法实现的遗憾界与理论下界匹配至对数因子。模拟结果表明,我们的算法实现了低遗憾,并且随着臂数量和更新成本的变化保持稳定。

英文摘要

Systems operating in dynamic environments require timely updates to sustain performance. For resource-intensive systems such as machine learning models and digital twins, strategically timing updates is essential. Updating too frequently wastes resources, while updating too infrequently leads to costly performance degradation. The problem is particularly challenging when the system's degradation pattern is unknown a priori, as is common in new operating environments. We formalize this challenge as a novel \emph{timing bandit} problem, where each arm represents a candidate update interval with a fixed update cost and an unknown, stochastic degradation cost. Three structural properties distinguish this setting from standard multi-armed bandits: selecting an interval commits the learner to multiple time slots before the next update; arm costs are composed of per-step degradation costs and a fixed update cost; and selecting a longer interval naturally reveals degradation at every intermediate step, providing consecutive feedback relevant to shorter intervals. By exploiting these structures, we develop Balanced Consecutive Arm Elimination (BCAE). BCAE achieves $\tilde{O}(\sqrt{T})$ regret, improving upon the $\tildeΩ(K\sqrt{T})$ regret of standard bandit algorithms in this setting, where $K$ is the number of candidate update intervals. We further propose an Optimism-Enhanced variant (OE-BCAE) that integrates lower-confidence-bound principles to improve empirical adaptivity while preserving the same regret order. Moreover, the regret bound achieved by our algorithms matches the theoretical lower bound up to logarithmic factors. Simulation results demonstrate that our algorithms achieve low regret and remain stable as both the number of arms and the update cost vary.

发表机构

  • City University of Hong Kong(香港城市大学)
  • The Chinese University of Hong Kong (Shenzhen)(香港中文大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

↑