arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29833math.OCstat.MEstat.ML

自适应马尔可夫随机逼近的收缩管集中界

Shrinking-Tube Concentration for Adaptive Markovian Stochastic Approximation

  • The University of Hong Kong(香港大学)
  • The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Jin Li, Ye Luo, Xiaowei Zhang

AI总结:

本研究为自适应马尔可夫随机逼近建立收缩管集中界,揭示容差收缩与逃逸概率衰减的权衡,并应用于库存学习,量化梯度精度对策略可靠性的影响。

AI中文摘要:

自适应算法在重塑生成其未来数据的动力学的同时,越来越多地做出决策。我们为受自适应马尔可夫链驱动的投影随机逼近建立了一个收缩管集中界。该界以高概率保证,在选定时间之后的每次迭代都保持在目标周围的容差内,且该容差随时间收紧。选定时间后任何逃逸的概率具有多项式衰减上界,而匹配的下界表明,在有限二阶矩条件下,其多项式指数通常无法改进。因此,该结果揭示了容差收缩速度与未来任何逃逸概率下降速度之间的尖锐权衡。我们还将分析扩展到具有额外鞅差噪声和可预测偏差的递推,表明鞅差噪声尺度的增长如何减慢逃逸概率界的衰减,而可预测偏差则限制了允许的管收缩。证明结合了后向核替换、有限时间均方误差界和分块最大首达逃逸论证。我们将该理论应用于具有缺货依赖需求和固定缺货成本的库存学习,并量化数值梯度精度对所得策略的全未来可靠性的影响。

英文摘要:

Adaptive algorithms increasingly make decisions while reshaping the dynamics that generate their future data. We establish a shrinking-tube concentration bound for projected stochastic approximation driven by an adaptive Markov chain. The bound guarantees, with high probability, that every iterate after a chosen time remains within a tolerance around the target that tightens over time. The probability of any exit after the chosen time admits a polynomially decaying upper bound, and a matching lower bound shows that its polynomial exponent cannot be improved in general under finite second moments. The result therefore identifies a sharp tradeoff between how quickly the tolerance shrinks and how rapidly the probability of any future exit decreases. We also extend the analysis to recursions with additional martingale-difference noise and predictable bias, showing how growth in the martingale-difference noise scale slows the decay of the exit-probability bound while predictable bias restricts the admissible tube shrinkage. The proof combines backward kernel replacement, a finite-time mean-squared-error bound, and a blockwise maximal first-exit argument. We apply the theory to inventory learning with stockout-dependent demand and fixed stockout costs, and quantify how numerical gradient accuracy affects the all-future reliability of the resulting policies.

↑