arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

折扣最小二乘的时间一致自归一化集中性:极限与修正

Time-Uniform Self-Normalized Concentration for Discounted Least Squares: Limits and Corrections

Yi-Shan Wu

arXiv 2608.19643首次发表:更新:

发表机构

Research Center for Information Technology Innovation, Academia Sinica(中央研究院资讯科技创新研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对折扣最小二乘的时间一致自归一化集中性,指出现有加权扩展的证明错误,修正了有限与无限时间的边界,为相关强化学习等下游分析提供了理论基础。

AI 中文摘要

自归一化集中不等式是多臂老虎机与强化学习分析中的标准工具。一种广泛使用的加权扩展声称,针对非平稳问题中的折扣最小二乘估计器存在类似的时间一致保证。一个带固定参数的简单标量高斯反例表明,所声称的有界半径会以概率1被突破。对于固定的折扣参数与正则化参数,我们进一步证明:当δ≤1/2且T/δ足够大时,在所述条件次高斯模型类上一致有效的任何确定性任意时间边界,在时间T之前的某个时刻必须至少达到R√log(T/δ)量级;对于非递减边界,该量级需在时间T处达到。我们指出了证明错误:不同的终止时刻使用不同的高斯混合分布,因此固定时间的混合分布不构成一个上鞅,而停时论证也无法修正这一缺陷。最后,我们证明该加权不等式在每个固定确定性时刻仍然有效,给出了有限时间与无限时间的有效修正,并讨论了其对下游分析的影响。

英文摘要

Self-normalized concentration inequalities are standard tools in bandit and reinforcement-learning analyses. A widely used weighted extension claims an analogous time-uniform guarantee for discounted least-squares estimators in non-stationary problems. A simple scalar Gaussian counterexample with a fixed parameter shows that the claimed bounded radius is crossed with probability one. For fixed discount and regularization parameters, we further show that, when $δ\leq1/2$ and $T/δ$ is sufficiently large, any deterministic anytime boundary valid uniformly over the stated conditionally sub-Gaussian model class must be at least of order $R\sqrt{\log(T/δ)}$ at some time by horizon $T$; for nondecreasing boundaries, this order is required at time $T$. We identify the proof error: different terminal times use different Gaussian mixing distributions, so the fixed-time mixtures do not form one supermartingale, and the stopping-time argument does not repair this failure. Finally, we show that the weighted inequality remains valid at each fixed deterministic time, give valid finite- and infinite-horizon corrections, and discuss consequences for downstream analyses.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑