通过策略记忆反向传播何时重要?物理信用、优化器更新与可观测性
When Does Backpropagating Through Policy Memory Matter? Physical Credit, Optimizer Updates, and Observability
浏览论文内容
中文总结 AI 辅助
本文研究策略记忆反向传播中切断存储路径的影响,发现物理信用、优化器更新和可观测性决定其重要性,建议用初始化训练和优化器更新评估记忆切断成本。
中文摘要 AI 辅助
具有记忆的策略可以沿两条反向路径学习:一条通过其动作产生的物理状态,另一条通过其存储的表示。Transformer-XL和截断式时间反向传播在存储历史处切断第二条路径,同时保留其值。我们研究这种切断何时重要。保持前向计算固定,仅改变导数边,我们在Transformer船舶轨迹模型和四旋翼跟踪策略中测量参数梯度、优化器应用的更新以及持续训练。在船舶模型中,当梯度流过所有早期物理状态时,分离键值缓存将梯度范数缩小到约十分之一,且旋转很小,但在一步物理信用下几乎不改变它。在这种强裁剪情况下,优化器而非梯度决定了更新差异的大小:全局范数裁剪消除了记忆切断图之间的大部分梯度差异,而AdamW将切断位置切换步骤中两个切断位置之间2%的梯度差异转化为高达31%的更新差异。在从初始化训练且具有0.20米/秒速度噪声的四旋翼中,移除记忆使跟踪误差增加43%,切断记忆梯度使其增加32%;在低噪声下,切断的平均成本超过记忆的价值。两步截断段没有带来可测量的收益,尽管在隐藏速度下,两步窗口捕获了大部分记忆价值;八步段消除了成本的一半到四分之三。仅在训练最后五分之一期间切换切断,在0.20-0.30米/秒速度下将其成本低估了约三倍,但在低噪声或隐藏速度下则不然。这些结果表明,通过从初始化开始训练来测量记忆切断的成本,并通过优化器应用的更新而非原始梯度来比较反向图。
英文摘要
Policies with memory can learn along two backward paths: through the physical states their actions produce and through the representations they store. Transformer-XL and truncated backpropagation through time cut the second path at stored history while keeping its values. We ask when this cut matters. Holding the forward computation fixed and varying only derivative edges, we measure parameter gradients, the updates the optimizer applies, and continued training in a Transformer vessel-trajectory model and a quadrotor tracking policy. In the vessel model, detaching the key-value cache shrank the gradient to about a tenth of its norm, with little rotation, when gradients flowed through all earlier physical states, but barely changed it under one-step physical credit. In this strongly clipped regime the optimizer, not the gradient, set how far updates differed: global-norm clipping removed most of the gradient difference between memory-cut graphs, whereas AdamW turned a 2% gradient difference between two placements of the cut into update differences of up to 31% at the step where the placement was switched. In a quadrotor trained from initialization with 0.20 m/s velocity noise, removing memory raised tracking error by 43% and cutting memory gradients raised it by 32%; at low noise the cut's mean cost exceeded the value of memory. Two-step truncation segments gave no measurable gain, although with hidden velocity a two-step window captured most of the value of memory; eight-step segments removed half to three quarters of the cost. Switching the cut on only for the last fifth of training understated its cost about threefold at 0.20-0.30 m/s, but not at low noise or with hidden velocity. These results suggest measuring the cost of a memory cut by training with it from initialization, and comparing backward graphs by the updates the optimizer applies rather than by raw gradients.