变点感知世界模型:在基于模型的强化学习中检测动力学变化并通过遗忘陈旧回放恢复
Changepoint-Aware World Models: Detecting Dynamics Shifts and Recovering by Forgetting Stale Replay in Model-Based RL
浏览论文内容
中文总结 AI 辅助
针对机器人动力学突变导致模型失效的问题,提出变点感知世界模型(CAWM),用在线CUSUM检验检测突变并遗忘陈旧回放,显著加快恢复速度,提升回报。
中文摘要 AI 辅助
机器人对其自身动力学的学习模型仅在动力学发生变化之前有效:执行器磨损、载荷转移、关节僵硬。一个持续训练却仿佛什么都没发生的基于模型的智能体适应缓慢,被充满陈旧经验的回放缓冲区拖累。我们提出了变点感知世界模型(CAWM),一个DreamerV3智能体,它利用自身的内部预测误差检测突然的动力学变化,使用在线CUSUM检验并对照滚动基线,该检验仅在突然变化而非缓慢学习漂移时触发。随后它遗忘陈旧的回放,保留学习到的表征同时清除过时数据。在两种与机器人相关的变化下的模拟运动中,即重力加倍和执行器增益减半,CAWM比被动重训练恢复得快得多。它还击败了一个强大的基线,该基线在检测到变化时重新生成新的动力学模型,即深度世界模型版的模型库方法。在变化发生时触发响应,CAWM在变化后前30k帧内相对于三个种子获得了+95到+153的回报提升,同时渐近性能与重新生成方法相当。闭环运行检测器在重力变化上重现了这一增益。该优势在两种变化类型中均成立,并且在变化足够严重以致旧数据确实过时时最为显著。
英文摘要
A robot's learned model of its own dynamics is only valid until those dynamics change: actuators wear, payloads shift, and joints stiffen. A model-based agent that keeps training as if nothing happened adapts slowly, dragged back by a replay buffer full of stale experience. We present Changepoint-Aware World Models (CAWM), a DreamerV3 agent that detects an abrupt dynamics shift from its own internal prediction error, using an online CUSUM test against a rolling baseline that fires only on abrupt change rather than on slow learning drift. It then forgets stale replay, keeping the learned representation while flushing obsolete data. On simulated locomotion under two robot-relevant shifts, doubled gravity and halved actuator gain, CAWM recovers substantially faster than passive retraining. It also beats a strong baseline that respawns a fresh dynamics model on detection, the deep-world-model analogue of model-bank methods. With the response triggered at the shift, CAWM gains +95 to +153 return in the first 30k post-shift frames over three seeds, while matching that respawn at asymptote. Running the detector in closed loop reproduces this gain on the gravity shift. The benefit holds across both shift types, and is largest when the shift is severe enough that old data is genuinely obsolete.