arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自适应多时间尺度强化学习

Adaptive Multi-Horizon Reinforcement Learning

Manoosh Samiei, Doina Precup, Paul Masset

arXiv 2607.20656首次发表:更新:

AI 中文总结

研究在复杂环境中强化学习的决策问题,提出自适应多时间尺度方法,能自适应选择和组合时间范围,无需手动调折扣因子,实验证明该方法可提高参数效率及增强适应性。

AI 中文摘要

在复杂多变的环境中进行有效决策,需要平衡短期和长期后果。在强化学习中,这种权衡通常通过固定折扣因子来控制,该因子施加单一指数折扣时间范围。然而,生物智能体表现出灵活和自适应的时间折扣,表明有效规划需要多个时间尺度。我们提出一种多时间尺度方法,自适应选择和组合时间范围,无需手动调整折扣因子就能稳健适应奖励结构变化。这种灵活性使该方法特别适用于涉及任务切换和不同环境配置的持续学习场景。实验表明,我们的方法能在一系列MiniGrid环境中识别有效折扣因子,包括由三个顺序变化任务组成的持续设置。这些结果表明,自适应时间折扣可提高参数效率,并增强人工和受生物启发学习系统的适应性。

英文摘要

Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement learning (RL), this trade-off is typically controlled through a fixed discount factor, which imposes a single exponentially discounted temporal horizon. However, biological agents exhibit flexible and adaptive temporal discounting, suggesting that effective planning requires multiple timescales. Here, we propose a multi-horizon approach that adaptively selects and combines temporal horizons, enabling robust adaptation to changes in reward structure without manual discount-factor tuning. This flexibility makes the method particularly suitable for continual learning scenarios involving task switches and varying environmental configurations. Empirically, we demonstrate that our approach identifies effective discount factors across a range of MiniGrid environments, including continual settings composed of three sequentially changing tasks. These results suggest that adaptive temporal discounting can improve parameter efficiency and enhance adaptability in both artificial and biologically inspired learning systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑