基于模型的强化学习中动态变化下的回放保留特性刻画
Characterizing Replay Retention Under Dynamics Shift in Model-Based Reinforcement Learning
查看机构详情
- Brown University(布朗大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究刻画了基于模型的强化学习中动力学变化下回放保留的权衡,提出变化幅度和年龄-陈旧度AUC两个量,并验证了回放策略选择依赖于变化幅度与动力学演化方式。
中文摘要 AI 辅助
适应机器人动力学变化需要从新数据中学习,同时不能丢弃可能仍然有用的经验。在持续基于模型的强化学习(RL)中,在动力学变化之前收集的回放数据可能会减慢适应速度,而移除这些数据则不必要地减少了可用的训练数据,并且如果早期动力学再次出现,移除数据可能代价高昂。我们研究了何时近期转换比完整回放历史更受青睐。两个量刻画了这一权衡:变化幅度和年龄-陈旧度曲线下面积(AUC),后者衡量转换年龄在多大程度上将陈旧数据与新鲜数据区分开来。遗忘陈旧数据在大的永久性变化后有所帮助,但在动力学重复出现且较旧数据再次变得有用时则有害。因此,选择回放策略取决于预测较旧数据何时会有帮助或有害。我们在两种运动形态、两种基于模型的RL算法以及真实世界RL基准扰动上测试了这些效应。由于在部署的机器人上无法获得真实陈旧度标签,我们评估了基于交互数据构建的估计器是否仍能提供在永久变化后选择回放策略所需的量。我们的结果表明,回放保留取决于变化幅度以及动力学的演化方式。
英文摘要
Adapting to changes in robot dynamics requires learning from new data without discarding experience that may still be useful. In continual model-based reinforcement learning (RL), replay collected before a dynamics change can slow adaptation, while removing it unnecessarily reduces available training data and can be especially costly if earlier dynamics return. We study when recent transitions are preferable to the full replay history. Two quantities characterize this trade-off: change magnitude and age-staleness area under the curve (AUC), measuring how well transition age separates stale from fresh data. Forgetting stale data helps after large permanent shifts but hurts when dynamics recur and older data becomes useful again. Choosing a replay strategy therefore depends on predicting when older data will help or hurt. We test these effects across two locomotion morphologies, two model-based RL algorithms, and Real-World RL benchmark perturbations. Because ground-truth staleness labels are unavailable on deployed robots, we evaluate whether an estimator built from interaction data can still provide the quantities needed to choose a replay strategy after permanent changes. Our results show that replay retention depends on change magnitude and on how the dynamics evolve.