Aftab:并行Q网络中CNN编码器与高级价值函数的综合基准
A Progressive Design Study of Visual Encoders and Value Estimation for Replay-Free Parallelized Q-Learning
浏览论文内容
中文总结 AI 辅助
本研究针对并行Q网络,设计评估8种CNN拓扑并集成多种Q学习扩展,提出复合架构Aftab,在Atari-57与Procgen Hard基准上均优于基线,已开源。
中文摘要 AI 辅助
深度强化学习的近期进展愈发倾向于简化、高度并行化的范式,值得注意的是,并行Q网络(Parallelized Q-Network, PQN)算法可实现稳定的离策略学习,且无需依赖计算成本高昂的经验回放缓冲区或目标网络。然而,在这些无缓冲区设置下运行的视觉编码器的表征能力与参数效率仍未得到充分探索。本研究系统探究了适用于PQN的卷积神经网络(Convolutional Neural Networks, CNN)的架构设计空间,设计并严格评估了8种不同的CNN拓扑结构,在严格的参数约束下优化样本效率;此外,通过集成哈达玛编码范式(Hadamax encoding paradigm)及分布式、集成式、对决式(dueling heads)等高级Q学习扩展方法,研究了表征与价值估计增强的影响。在Atari-57基准上开展的大量实验表明,本研究提出的复合架构Aftab实现了6.479的四分位均值(Interquartile Mean, IQM)人类归一化得分,较标准PQN基线建立了0.86的改进概率;此外,在高度非平稳的Procgen Hard基准上进行的结构鲁棒性评估证实了分布外泛化能力,Aftab的IQM Procgen归一化得分为0.418,而基线为0.382。最终,本研究为无模型强化学习建立了高效、概率上更优的结构参考,同时保留了无缓冲区、并行化优化的简洁性与内存效率。完整的Aftab框架,包括所有模型定义、训练配置及原始实验日志,已开源并可在GitHub仓库获取:this https URL
英文摘要
Replay-free parallelized Q-learning removes the large experience replay buffers and target networks used by conventional deep Q-learning, but the role of network architecture in this training regime remains comparatively underexplored. We investigate this question through a progressive three-phase study within the Parallelized Q-Network (PQN) framework. First, we compare eight convolutional encoder topologies on Atari-57 under a common training protocol while jointly considering performance and computational complexity. Second, we integrate Hadamax-style multiplicative feature interactions and explicit pooling into the selected encoder hierarchy. Third, with the visual representation fixed, we compare complete categorical-dueling, ensemble-dueling, and categorical ensemble-dueling value-estimation configurations. The resulting architecture, Aftab, achieves an interquartile mean human-normalized score of $6.592$ on Atari-57, compared with $2.715$ for our independently rerun PQN reference, with a game-level Probability of Improvement of $0.86$. After completing all architecture selection on Atari-57, we evaluate Aftab on Procgen Hard. Aftab achieves a terminal IQM normalized score of $0.418$ compared with $0.382$ for PQN and increases the normalized area under the learning curve from $0.216$ to $0.541$, although terminal performance remains heterogeneous across environments. These results show that visual topology, multiplicative representation, and downstream value-estimation design can substantially affect replay-free Q-learning, and that their benefits should be evaluated jointly with computational complexity. The complete Aftab framework, including model definitions, training configurations, reproducibility settings, and raw experimental logs, is open-sourced at https://github.com/tahashieenavaz/aftab
发表机构
- University of Padua(帕多瓦大学)
机构由 AI 辅助整理,请以论文原文为准。