数据驱动的布朗反射控制
Data-Driven Brownian Reflection Control
浏览论文内容
中文总结 AI 辅助
针对漂移和波动率未知的布朗反射控制问题,提出先学习后优化、自适应更新及全历史自适应更新三种算法,分别取得不同阶的悔值界,并从理论上分解分析悔值构成。
中文摘要 AI 辅助
我们研究了漂移和波动率未知的布朗模型的数据驱动反射控制问题。首先提出一种先学习后优化(LTO)算法:在探索阶段估计与策略相关的参数,将估计值代入最优性方程,再利用所得策略,实现了$O(\sqrt{T})$的有限时间期望悔值界。我们进一步提出两种算法:自适应更新(AU)和全历史自适应更新(AU-FH),它们持续更新估计器和反射水平,达到了更优的$O(\log T)$悔值界。值得注意的是,AU-FH算法利用所有历史数据,在数值模拟中表现更优。我们的分析将悔值分解为探索、瞬态和学习三个部分。由非平稳性导致的瞬态悔值,可通过转移半群相对于平稳性的时间积分偏差(在持有成本上的取值)来界定,该偏差还可通过受控反射布朗运动(RBM)指数收敛的Foster-Lyapunov不等式进一步界定。对于学习悔值,我们建立了估计器的局部正则性性质、一致性以及均方误差界,这些性质控制了学习到的反射策略与最优反射策略之间的平稳成本差距。此外,我们利用移动边界Skorokhod映射的单调性,通过与固定边界RBM的路径比较,推导了AU算法切换状态的矩界。
英文摘要
We study a data-driven reflection control problem for a Brownian model with unknown drift and volatility. We first propose a learn-then-optimize (LTO) algorithm: it estimates the policy-relevant parameter during exploration, plugs the estimate into the optimality equation, and exploits the resulting policy---achieving an $O(\sqrt{T})$ finite-time expected regret bound. We further propose two algorithms, adaptive-updating (AU) and full-history adaptive-updating (AU-FH), which continuously update the estimator and reflecting level, attaining an improved $O(\log T)$ regret bound. Notably, AU-FH algorithms leverages all historical data, yielding better performance in numerical simulations. Our analysis decomposes regret into exploration, transient, and learning components. Transient regret from nonstationarity is bounded by the time-integrated deviation of the transition semigroup from stationarity evaluated on the holding cost, which can be further bounded via a Foster-Lyapunov inequality for exponential convergence of the controlled reflected Brownian motion (RBM). For learning regret, we establish local regularity properties together with consistency and mean-squared error bounds for the estimator, which control the stationary cost gap between the learned and optimal reflection policies. In addition, we leverage the monotonicity of the moving boundary Skorokhod map to derive moment bounds for the AU algorithms' switching states, via pathwise comparison with fixed-boundary RBMs.
发表机构
- Rice University(莱斯大学)
- Academy of Mathematics and Systems Science, Chinese Academy of Sciences(中国科学院数学与系统科学研究院)
- University of Chinese Academy of Sciences(中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。