发表机构
Donghua University; Westlake University; Imperial College London; East China University of Science and Technology(东华大学; 西湖大学; 帝国理工学院; 华东理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SpikeCredit框架,利用脉冲神经网络的时间信用载体机制,在稀疏奖励强化学习中通过双通路设计实现高效信用分配,显著提升多个MuJoCo任务的性能。
AI 中文摘要
稀疏奖励下的强化学习(RL)具有挑战性,因为延迟的结果对于哪些中间计算导致了成功或失败提供的指导很少。我们认为,可靠的信用分配需要策略动态,这些动态随时间保留并暴露与信用相关的信息,我们将这一角色形式化为时间信用载体(TCCs),而脉冲神经网络(SNNs)通过分级膜电位轨迹和事件驱动的脉冲自然地实现了这一角色。基于这一假设,我们提出了SpikeCredit,一个基于SNN的稀疏奖励RL框架,该框架首先执行任务自适应的TCC选择,然后在一个快速的TCC读取通路(其中自运动反馈约束使用基于本地行为线索来约束转移级信用恢复)和一个慢速的TCC写入通路(其中信用目标轨迹对齐将恢复的信用反馈给演员,使未来的TCC动态更具信用可读性)之间闭环。在稀疏奖励的MuJoCo任务中,SpikeCredit在Ant、Hopper、Swimmer和Walker2d上分别将Last10回报相对于稀疏SNN基线提高了+1169%、+953%、+723%和+1781%,并在Swimmer上超过密集奖励基线+113%。机制分析进一步表明,与稀疏SNN基线相比,与密集奖励的对齐更强。这些结果将脉冲动态定位为稀疏奖励RL的信用保留基质。
英文摘要
Reinforcement learning (RL) with sparse rewards is challenging because delayed outcomes provide little guidance about which intermediate computations caused success or failure. We argue that reliable credit assignment requires policy dynamics that preserve and expose credit-relevant information over time, a role we formalize as Temporal Credit Carriers (TCCs) and that spiking neural networks (SNNs) naturally fulfill through graded membrane traces and event-driven spikes. Based on this hypothesis, we propose SpikeCredit, an SNN-based framework for RL with sparse rewards that first performs task-adaptive TCC selection and then closes the loop between a fast TCC-reading pathway, where self-motion feedback constraint uses local behavior-grounded cues to constrain transition-level credit recovery, and a slow TCC-writing pathway, where credit-targeted trace alignment feeds recovered credit back into the actor to make future TCC dynamics more credit-readable. Across sparse-reward MuJoCo tasks, SpikeCredit improves Last10 return over sparse SNN baselines by +1169% on Ant, +953% on Hopper, +723% on Swimmer, and +1781% on Walker2d, and exceeds the dense-reward baseline on Swimmer by +113%. Mechanistic analyses further show substantially stronger alignment with dense rewards than the sparse SNN baseline. These results position spiking dynamics as credit-preserving substrates for sparse-reward RL.
Comments14 pages, 9 figures