arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31586cs.LG

信任引导的决策变换器

Trust Guided Decision Transformer

  • International Institute of Information Technology Bangalore(国际信息技术学院班加罗尔)
  • IBM Research Bangalore(IBM研究院班加罗尔)

机构由 AI 辅助整理,请以论文原文为准。

Chainesh Gautam, Raghuram Bharadwaj Diddigi, Chandramouli Kamanchi, Pankaj Dayama, Sumanta Mukherjee, Kameshwaran Sampath

AI总结:

针对决策变换器在长程轨迹中因上下文漂移导致的性能下降,提出信任引导的决策变换器(TGDT),通过预测误差校准选择可信上下文,再用评论家选动作,在D4RL任务上提升回报并减少持续高误差。

AI中文摘要:

决策变换器(Decision Transformer)在长程轨迹上的性能会下降,因为其条件上下文漂移出训练分布。我们表明,这种漂移可以通过模型自身的下一状态预测误差来观察,该误差在轨迹展开过程中上升并持续保持高位,从而直接指示上下文何时变得不可靠。我们引入了信任引导的决策变换器(Trust Guided Decision Transformer, TGDT),它在应用价值引导之前先选择上下文。在每一步中,TGDT 使用滚动下一状态预测误差评估几个最近的上下文后缀,并通过分裂共形预测(split conformal prediction)针对离线数据进行校准。它仅保留误差保持在校准阈值内的后缀,然后使用冻结的评论家(frozen critic)在受信任的后缀中选择最高价值的动作。这逆转了仅基于价值的弹性选择(value only elastic selection)的顺序,在后一种方法中,评论家可能选择一个由模型自身标记为不可靠的上下文生成的动作。在 D4RL 导航和运动任务上的实验表明,状态预测、评论家引导和硬上下文重置各自只能解决部分问题。TGDT 减少了持续高误差的运行,并提高了相对于原始决策变换器、基于重置的上下文控制和仅基于价值的上下文选择的回报。

英文摘要:

Decision Transformer performance degrades on long rollouts because the conditioning context drifts out of the training distribution. We show that this drift is visible through the model's own next state prediction error, which rises during rollout and stays elevated, giving a direct signal of when context has become unreliable. We introduce Trust Guided Decision Transformer (TGDT), which selects context before applying value guidance. At each step, TGDT evaluates several recent context suffixes using rolling next state prediction error, calibrated against held out offline data via split conformal prediction. It keeps only suffixes whose error stays within the calibrated threshold, then uses a frozen critic to choose the highest value action among the trusted suffixes. This reverses the order used by value only elastic selection, where the critic may choose an action generated from a context the model itself has flagged as unreliable. Experiments on D4RL navigation and locomotion tasks show that state prediction, critic guidance, and hard context reset each solve only part of the problem. TGDT reduces persistent high error runs and improves return over vanilla Decision Transformer, reset based context control, and value only context selection.

补充信息

↑