发表机构
CISPA Helmholtz Center for Information Security; Technical University of Munich(CISPA亥姆霍兹信息安全中心; 慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一种基于神经ODE的正则化方法,将其集成到Actor-Critic算法中,在A2C的Atari基准和PPO的网格世界环境中显著提升了强化学习智能体的性能。
AI 中文摘要
应用于序列决策任务的神经网络通常依赖于环境状态的隐表征。环境动态决定了语义状态的演化,而对应的隐状态转移通常是隐式的,这可能导致两者之间出现错位。我们通过将马尔可夫决策过程(MDP)轨迹与常微分方程(ODE)流进行类比,显式建模隐动态:两者均满足当前状态完全决定后续状态。基于此观点,我们提出一种基于神经ODE的正则化方法,强制隐嵌入遵循一致的ODE流,从而使表征学习与环境动态对齐。该方法虽广泛适用于深度学习智能体,但我们通过将其集成到Actor-Critic算法中,验证了其在强化学习中的有效性。在A2C的各类标准Atari基准测试及PPO的网格世界环境中,我们的方法均取得了显著的性能提升。
英文摘要
Neural networks applied to sequential decision-making tasks typically rely on latent representations of environment states. While environment dynamics dictate how semantic states evolve, the corresponding latent transitions are usually left implicit, creating a potential misalignment between the two. We propose to model latent dynamics explicitly by drawing an analogy between Markov decision process (MDP) trajectories and ordinary differential equation (ODE) flows: in both cases, the current state fully determines its successors. Building on this view, we introduce a neural ODE-based regularization method that enforces latent embeddings to follow consistent ODE flows, thereby aligning representation learning with environment dynamics. Although broadly applicable to deep learning agents, we demonstrate its effectiveness in reinforcement learning by integrating it into Actor-Critic algorithms. Our approach yields major performance gains across various standard Atari benchmarks for A2C and gridworld environments for PPO.