arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19397cs.LG

内存合并深度Q网络:用于稳定值学习的灵敏度加权目标更新

Memory Merge DQN: Sensitivity Weighted Target Updates for Stable Value Learning

  • Deakin University(迪肯大学)
  • Federation University(联邦大学)
  • UNSW(新南威尔士大学)

机构由 AI 辅助整理,请以论文原文为准。

Adrian Ly, Richard Dazeley, Peter Vamplew, Sunil Aryal, Francisco Cruz

中文总结 AI 辅助

研究针对深度Q网络目标网络更新的权衡问题,提出内存合并深度Q网络机制,通过保留近期在线网络副本记忆,依Q值灵敏度合并参数构建目标网络,实验表明该方法能提升稳定性与最终性能。

中文摘要 AI 辅助

深度Q网络使用目标网络来稳定自举值学习,但标准的硬拷贝更新也带来了权衡。固定目标网络可提高短期稳定性,但每次硬更新都会突然用最新的在线网络替换目标参数并丢弃近期参数历史。这可能导致引导目标突然变化,并可能去除训练后期仍有用的值函数结构。本文介绍了内存合并深度Q网络,一种目标网络更新机制,它保留近期在线网络副本的短记忆,并根据Q值灵敏度合并网络参数来构建目标网络,而非仅复制最新在线网络。该方法受Fisher权重模型合并启发,但使用Q值灵敏度而非Fisher信息作为加权信号。本文在Atari环境中对内存合并深度Q网络与深度Q网络、平均深度Q网络、带层归一化的深度Q网络和PQN(带梯度裁剪)进行了评估。结果表明,内存合并深度Q网络具有高度竞争力,在评估方法中获得第一名最终性能结果的数量最多,击败了深度Q网络、平均深度Q网络和PQN(带梯度裁剪),并在保留有用值函数参数有益的几个游戏中取得了显著收益。这些发现表明,选择性合并近期参数权重和历史可以提高深度Q网络智能体的稳定性和最终性能,并且目标网络设计是在长期值学习中保留有用值函数结构的重要机制。

英文摘要

Deep Q-networks use target networks to stabilise bootstrapped value learning, but the standard hard copy update also introduces a tradeoff. Holding the target network fixed, improves short term stability, yet each hard update abruptly replaces the target parameters with the newest online network and discards recent parameter history. This can produce sudden changes in the bootstrap target and may remove value function structure that remains useful later in training. This paper introduces Memory Merge DQN, a target network update mechanism that maintains a short memory of recent historical online network copies and constructs the target network by merging network parameters based on the Q-value sensitivity rather than copying only the newest online network. Memory Merge gives greater influence to parameters that remain locally important for current Q-value behaviour, while using a recency prior to keep the merged target close to the latest online parameters. The method is inspired by Fisher Weight Model Merging, but uses Q-value sensitivity rather than Fisher information as the weighting signal. This paper evaluates Memory Merge DQN on Atari environments against DQN, Averaged DQN, DQN with layer normalisation, and PQN (with gradient clipping). The results show that Memory Merge DQN is highly competitive and it achieves the largest number of first place final performance results among the evaluated methods, beats DQN, Averaged DQN, and PQN (with gradient clipping), and produces substantial gains in several games where preserving useful value-function parameters appears beneficial. These findings suggest that selectively merging recent parameter weights and history can improve the stability and final performance of DQN agents, and that target network design is an important mechanism for preserving useful value function structure during long horizon value learning.

↑