arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

决策泰坦:离线强化学习中的测试时训练用于长期记忆

Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning

Jude Waide, Robert Lieck

arXiv 2610.01513首次发表:更新:

发表机构

University of Durham(杜伦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将测试时训练框架应用于离线强化学习,提出决策泰坦模型,在X-Maze环境中验证其能学习超出上下文窗口20倍的长期依赖,并揭示时间嵌入与信息编码对泛化能力的关键影响。

AI 中文摘要

长期依赖仍然是人工智能领域顺序决策中的一个主要挑战:RNN遭受梯度消失和基于向量的隐藏状态表达能力有限的问题,而基于Transformer的模型则受到注意力二次方扩展的限制。最近的工作提出利用测试时训练(TTT)框架来解决这个问题,该框架通过在训练和测试时进行梯度下降,将情景记忆存储在神经网络的参数中。这种方法在自然语言处理领域已取得成功,然而,据我们所知,它尚未应用于强化学习(RL)领域,也没有研究分析这种记忆实际如何运作。在本文中,我们通过将决策Transformer与TTT层结合(称为决策泰坦),研究TTT框架在离线RL中的潜力。我们在X-Maze环境中分析模型的性能和特性,X-Maze是T-Maze的扩展,旨在测试顺序记忆,并通过随时间可视化门控值来研究记忆机制的学习方式。我们的主要发现是,决策泰坦能够学习比上下文窗口长20倍范围的长期依赖,泛化到训练数据1.7倍的长度,但关键的是,时间泛化取决于所使用的时间嵌入,而学习长期依赖的能力取决于相关信息如何编码。

英文摘要

Long-term dependencies remain a major challenge for sequential decision-making in the field of AI: RNNs suffer from vanishing gradients and the limited expressivity of vector-based hidden states, whilst Transformer-based models are limited by the quadratic scaling of attention. Recent work has proposed tackling this problem with the Test-Time Training (TTT) framework, which stores episodic memories in the parameters of a neural network through gradient descent at both train and test-time. This approach has seen success in the domain of Natural Language Processing, however, to the best of our knowledge it has not yet been applied to the domain of Reinforcement Learning (RL), nor has there been a study analysing how this memory practically functions. In this paper, we study the potential of the TTT framework for offline RL by augmenting a Decision Transformer with TTT layers, dubbed the Decision Titan. We analyse performance and properties of the model in the X-Maze environment, an extension of T-Maze designed to test sequential memory, and investigate how the memory mechanism learns by visualising gate values over time. Our key findings are that Decision Titan can learn long-term dependencies with ranges 20x longer than the context window, generalises to lengths 1.7x the training data, but crucially temporal generalisation depends on the time embeddings used, and the ability to learn long-term dependencies depends on how the relevant information is encoded.

CommentsAccepted at ICML 2026 Workshop on Decision-Making from Offline Datasets to Online Adaptation: Black-Box Optimization to Reinforcement Learning

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑