arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

动态上下文调度:超越静态宇宙的学习

Dynamic Context Scheduling: Learning Beyond the Static Universe

Martin Mráz, André Biedenkapp

arXiv 2608.20799首次发表:更新:

发表机构

University of Freiburg(弗赖堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出动态上下文调度方法,构建DYNAMICCARLENV框架,在CartPole等任务中验证其在分布外场景的性能,且对复杂环境可提升分布内表现,自动搜索多阶段课程可高效发现优化泛化的调度方案。

AI 中文摘要

我们将动态上下文调度作为上下文强化学习的训练工具展开研究。我们不再将单回合内的上下文变化视为部署中的实际情况,而是将其作为一种可控的塑造机制,据此,上下文会根据预设的调度规则在每个训练回合内演变,让智能体策略接触到环境参数空间中更丰富、时间结构更清晰的区域。我们推出了DYNAMICCARLENV框架,该框架可通过可插拔的调度族(如正弦偏移或余弦退火)封装上下文环境。在带有CARL上下文的CartPole、BipedalWalker和VehicleRacing任务中,我们证明动态调度在分布外(OOD)场景下与静态上下文基线表现相当或更优;值得注意的是,对于更复杂的BipedalWalker和VehicleRacing环境,我们还实现了更高的分布内(ID)评估性能。初步研究结果表明,自动搜索多阶段课程能够成功发现可提升泛化能力的调度方案,其表现与对单阶段调度器进行的广泛网格搜索相当。

英文摘要

We study dynamic context scheduling as a training instrument for contextual re- inforcement learning. Rather than treating intra-episode context variation as a deployment reality, we treat it as a controlled shaping mechanism. Thereby, context evolves within each training episode according to a predetermined schedule, expos- ing the policy to a richer and more temporally structured region of the environment parameter space. We introduce DYNAMICCARLENV, a framework that wraps contextual environments with pluggable schedule families, such as sinusoidal off- sets or cosine annealing. Across CartPole, BipedalWalker and VehicleRacing with CARL contextualization, we show that dynamic schedules match or outperform static context baselines in the out-of-distribution (OOD) regimes. Interestingly, for the more complex BipedalWalker and VehicleRacing environments we also achieve higher in-distribution (ID) evaluation performance. Preliminary findings indicate that automatic search for multi-stage curricula can successfully discover schedules that improve generalization, performing comparably to extensive grid search over single-stage schedulers.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑