共模误差限制低时间步深度脉冲Q网络
Common-Mode Errors Limit Low-Timestep Deep Spiking Q-Networks
- Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对低时间步深度脉冲Q网络性能退化问题,通过误差分解发现共模误差是主因,提出CMC-DSQN用辅助ANN补偿共模误差,在Atari和MiniAtar上显著提升性能。
AI中文摘要:
脉冲神经网络(SNN)提供稀疏且事件驱动的计算,使其在边缘设备上的能量受限强化学习(RL)中具有吸引力。在基于价值的强化学习中,深度脉冲Q网络(DSQN)将这种效率与动作价值估计相结合用于决策。然而,现有的DSQN通常需要多个模拟时间步才能获得有竞争力的性能,增加了计算和能量成本,而减少时间步可能导致性能显著下降。我们从Q值估计误差的角度研究这种退化。通过将跨动作的误差分解为共模和差模分量,我们发现低时间步DSQN受到跨动作价值共享的共模误差的不成比例影响,这些误差通过自举目标对时间差分学习尤其有害。基于这一发现,我们提出了共模补偿深度脉冲Q网络(CMC-DSQN),它使用辅助人工神经网络(ANN)来补偿SNN输出中的共模误差。在推理时,可以直接从SNN输出执行贪婪动作选择,从而可以完全移除辅助ANN并保持SNN的能量效率。在Atari和MiniAtar环境上的大量实验证明了在低时间步设置下的显著性能改进。CMC-DSQN在T=2时优于最先进的DSQN基线近20%,并在T=4时进一步超过ANN基线。
英文摘要:
Spiking neural networks (SNNs) offer sparse and event-driven computation, making them attractive for energy-constrained reinforcement learning (RL) on edge devices. In value-based RL, deep spiking Q-networks (DSQNs) combine such efficiency with action-value estimation for decision making. However, existing DSQNs often require multiple simulation timesteps for competitive performance, increasing computational and energy costs, whereas reducing the timesteps can cause substantial performance degradation. We investigate this degradation from the perspective of Q-value estimation errors. By decomposing errors across actions into common-mode and differential-mode components, we find that low-timestep DSQNs suffer disproportionately from common-mode errors shared across action values, which are particularly detrimental to temporal-difference learning through bootstrapped targets. Based on this finding, we propose Common-Mode Compensation Deep Spiking Q-Network (CMC-DSQN), which uses an auxiliary ANN to compensate for common-mode errors in the SNN outputs. At inference, greedy action selection can be performed directly from the SNN outputs, allowing the auxiliary ANN to be completely removed and preserving the energy efficiency of SNNs. Extensive experiments on Atari and MiniAtar environments demonstrate substantial performance improvements under low-timestep settings. CMC-DSQN outperforms state-of-the-art DSQN baselines by nearly $20\%$ at $T=2$ and further surpasses the ANN baseline at $T=4$.