发表机构
Shanghai Jiao Tong University; Alibaba Group; University at Buffalo(上海交通大学; 阿里巴巴集团; 纽约州立大学布法罗分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出QTPT方法,用Q目标替代行为克隆进行上下文强化学习预训练,在理论和实验上证明其对弱数据更鲁棒,优于监督行为预测。
AI 中文摘要
现有的上下文强化学习方法主要使用监督行为预测目标来预训练Transformer。这能够从上下文中推断任务,但使得学习到的策略强烈依赖于离线动作的质量:当轨迹较弱或次优时,模仿本身成为有偏的学习信号。我们提出Q目标预训练Transformer(QTPT),该方法保留上下文条件Transformer架构,但将行为克隆替换为Bellman风格的Q目标目标。因此,QTPT学习利用上下文中的奖励和转移来估计动作价值,而非简单模仿行为策略。我们在随机线性老虎机和有限时域MDP中对QTPT进行理论分析,显示出相比监督预训练对数据质量更强的鲁棒性。实证上,在具有随机或次优数据的受控强化学习基准上,QTPT优于监督行为预测,并考察了在D4RL Kitchen和AntMaze上的扩展。补充实验评估了骨干网络鲁棒性、元强化学习比较、任务一致上下文以及不支持动作的价值高估。这些比较区分了Q目标预训练带来的益处与离线覆盖的剩余局限性。
英文摘要
Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. This enables task inference from context, but makes the learned policy strongly depend on the quality of offline actions: when trajectories are weak or suboptimal, imitation itself becomes a biased learning signal. We propose Q-Target Pretrained Transformers (QTPT), which keeps the context-conditioned Transformer architecture but replaces behavior cloning with a Bellman-style Q-target objective. QTPT therefore learns to use rewards and transitions in the context to estimate action values, rather than simply imitating the behavior policy. We theoretically analyze QTPT in stochastic linear bandits and finite-horizon MDPs, showing stronger robustness to data quality than supervised pretraining. Empirically, QTPT improves over supervised behavior prediction on controlled RL benchmarks with random or suboptimal data, and we examine extensions to D4RL Kitchen and AntMaze. Supplementary experiments evaluate backbone robustness, meta-RL comparisons, task-coherent context, and unsupported-action value overestimation. These comparisons distinguish the benefits of Q-target pretraining from the remaining limitations of offline coverage.
Comments41 pages, 6 figures. Accepted at NeurIPS 2026