arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23775cs.LGcs.SYeess.SY

Transformer编码器加速鲁棒强化学习的统计收敛性

Statistical Convergence of Transformer Encoder-Accelerated Robust Reinforcement Learning

Suman Banerjee, Hiroyasu Tsukamoto

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出一种Transformer热启动的鲁棒强化学习算法,采用R-污染模型和保形预测,通过预认证停止规则在扰动迷宫环境中显著减少初始误差并加速收敛,同时提供更紧密的误差界限。

中文摘要 AI 辅助

在马尔可夫决策过程中,获取最优动作价值函数在大状态-动作空间中计算密集。在本研究中,我们为一种由基于Transformer的动作价值函数预测进行热启动的鲁棒强化学习算法提供了统计上严格的收敛性结果,其中自然语言提示编码任务规范。我们的框架采用R-污染模型来刻画状态转移核中的不确定性,并利用保形预测通过由收缩贝尔曼残差构建的轨迹级非一致性分数来认证收敛性。所得的保形分位数同时限制了所有迭代中运行与最优动作价值函数之间的差距,从而产生一个预认证的停止规则,该规则对真实转移核的知识需求很少。在不同规模和污染水平的扰动迷宫环境上的数值案例研究证实,基于Transformer的热启动可显著减少初始误差并加速收敛,而所提出的保形界限比现有保证更紧密地跟踪真实误差轨迹。

英文摘要

Obtaining the optimal action-value function in Markov decision processes is computationally intensive in large state--action spaces. In this study, we present statistically rigorous convergence results for a robust reinforcement learning algorithm warm-started by a transformer-based action-value function prediction, where natural language prompts encode task specifications. Our framework adopts the R-contamination model to characterize uncertainty in the state transition kernel, and employs conformal prediction to certify convergence via trajectory-level nonconformity scores constructed from the contracting Bellman residual. The resulting conformal quantile bounds the gap between the running and optimal action-value functions simultaneously over all iterations, thereby yielding a pre-certified stopping rule that requires little knowledge of the true transition kernel. Numerical case studies on perturbed maze environments of varying size and contamination level confirm that the transformer-based warm start measurably reduces the initial error and accelerates convergence, while the proposed conformal bounds track the true error trajectory more tightly than existing guarantees.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

↑