rlpyt: A Research Code Base for Deep Reinforcement Learning in PyTorch
专题命中 仓库级理解 :repository(abstract);分类 cs.AI、cs.LG
Comments v2: Updated learning curves for SAC and TD3, improved by bootstrapping value-function when trajectory ends due to time limit, and switching to newer SAC version, now referenced