arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可塑性的贝尔曼最优性方程

A Bellman Optimality Equation for Plasticity

Jeremy Lucas, Doina Precup

arXiv 2609.10776首次发表:更新:

AI 中文总结

本文针对持续强化学习中的可塑性优化问题,基于Abel等人对可塑性的新定义,证明了在马尔可夫决策过程中存在可塑性的贝尔曼最优性方程,为可塑性优化提供了理论基础。

AI 中文摘要

在持续强化学习中,仔细管理稳定性-可塑性权衡仍然是一个核心挑战。Abel等人(2025)的最新工作通过将可塑性定义为从智能体观察到其动作的广义有向信息,并将赋权(empowerment)定义为从其动作到其观察的广义有向信息,从而形式化了这一困境。该公式成功地将传统的稳定性-可塑性权衡重新定义为赋权-可塑性权衡。然而,尽管存在大量关于优化赋权的文献,但目前尚无研究涉及在这种新定义下优化可塑性。本文展示了在马尔可夫决策过程中优化可塑性的初步工作。我们证明,存在一个类似于先前赋权工作的可塑性优化贝尔曼最优性方程。

英文摘要

In continual reinforcement learning, carefully managing the stability-plasticity tradeoff remains a core challenge. Recent work by Abel et al. (2025) formalized this dilemma by defining plasticity as the generalized directed information from an agent's observations to its actions, and empowerment as the generalized directed information from its actions to its observations. This formulation successfully reframes the traditional stability-plasticity tradeoff as an empowerment-plasticity tradeoff. However, while extensive literature exists on optimizing for empowerment, there is currently no research addressing the optimization of plasticity under this new definition. This paper presents preliminary work toward optimizing plasticity within Markov decision processes. We show that there exists a Bellman optimality equation for optimizing plasticity similar to previous work for empowerment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑