arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

非平稳性下的上下文强化学习:一项综述

In-Context Reinforcement Learning under Non-Stationarity: A Survey

A Run, Ziluo Ding

arXiv 2607.11906首次发表:更新:

发表机构

Beijing Innovation Center of Humanoid Robotics(北京人形机器人创新中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该综述围绕非平稳性下的上下文强化学习展开,探讨预训练或微调决策模型在交互中推断规则改善行为的能力。将其定义为策略固定时通过上下文适应问题,关联多种相关技术,并从变化内容、展开方式及智能体可观察性三方面组织文献。

AI 中文摘要

决策预训练变压器、算法蒸馏、长上下文元强化学习和检索增强智能体的发展,重新引发了人们对上下文强化学习(ICRL)的兴趣,即预训练或微调后的决策模型在不进行测试时参数更新的情况下,从交互上下文中推断潜在任务规则并改善未来行为的能力。现有ICRL综述主要围绕预训练目标、架构、上下文格式、评估协议和理论机制来组织该领域,而非平稳设置相对研究不足。在变化的环境中,积累的上下文不仅是关于固定任务的更多证据,奖励规范、转移核、观察通道、动作接口、约束模型或示范和记忆分布可能与当前状态不一致。因此,以前有用的上下文可能会过时、产生误导,或者在旧状态恢复时再次变得有用。我们将非平稳ICRL综述为在部署策略参数保持固定的情况下通过上下文进行适应的问题:策略必须推断当前决策规则以及其积累的证据中哪些部分仍然支持该规则。我们定义了非平稳ICRL,将其与元强化学习、决策序列建模、检索增强强化学习、值和模型感知ICRL以及奖励反馈智能体相关联,并沿着三个问题组织文献:什么发生了变化、变化如何展开以及智能体对变化的可观察程度如何。

英文摘要

The development of decision-pretrained transformers, algorithm distillation, long-context meta-RL, and retrieval-augmented agents has renewed interest in in-context reinforcement learning (ICRL): the ability of a pretrained or fine-tuned decision model to infer latent task rules and improve future behavior from interaction context, without test-time parameter updates. This line of work asks when trial-and-error evidence, rewards, transitions, demonstrations, feedback, or retrieved experience can make learning-like computation happen inside the context window. However, existing surveys of ICRL mainly organize the field around pretraining objectives, architectures, context formats, evaluation protocols, and theoretical mechanisms, while the non-stationary setting remains comparatively underexamined. In changing environments, accumulated context is not merely more evidence about a fixed task: the reward specification, transition kernel, observation channel, action interface, constraint model, or demonstration and memory distribution can fall out of alignment with the current regime. Previously useful context can therefore become stale, misleading, or useful again when an old regime returns. We survey non-stationary ICRL as the problem of adapting through context while deployed policy parameters remain fixed: the policy must infer both the current decision rule and which parts of its accumulated evidence still support that rule. We define non-stationary ICRL, relate it to meta-RL, decision sequence modeling, retrieval-augmented RL, value- and model-aware ICRL, and reward-feedback agents, and organize the literature along three questions: what changes, how the change unfolds, and how observable the change is to the agent.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑