AI 中文总结
本文提出Momba,将观测与特征归一化等三项神经网络设计进展与熵正则化MORL算法结合,在标准连续控制基准上大幅提升了MORL算法生成解集的质量。
AI 中文摘要
深度强化学习(RL)的近期进展表明,改进神经网络架构无需改变底层算法即可在样本效率和渐近性能上取得显著提升。相比之下,旨在发现一组平衡冲突目标间权衡关系的策略的多目标强化学习(MORL)研究,主要聚焦于算法创新,却忽视了架构领域的探索。尽管最优策略和价值函数会因权衡关系的不同而存在显著差异,但MORL算法通常以权衡条件下的简单前馈网络来表示它们,这引发了一个问题:使用更具表达力的函数近似器能否提升算法性能。本文将近期神经网络设计的进展:(i)观测与特征归一化、(ii)权重归一化、(iii)分布回报建模,与熵正则化MORL算法相结合。在标准连续控制基准上的实证结果表明,这些改动大幅提升了所生成解集的质量,且无需对底层算法进行重大修改。
英文摘要
Recent advances in deep reinforcement learning (RL) have shown that improving neural network architectures can yield substantial gains in sample efficiency and asymptotic performance without altering the underlying algorithms. In contrast, work on multi-objective reinforcement learning (MORL), which aims to discover a set of policies that balance trade-offs among conflicting objectives, has predominantly focused on algorithmic innovations, leaving the area of architectures underexplored. While the optimal policies and value functions can differ significantly depending on the trade-offs, MORL algorithms commonly represent them with simple feedforward networks conditioned on the trade-off. This raises the question of whether the performance of the algorithms could be improved with more expressive function approximators. In this paper, we integrate recent advances in neural network design: (i) observation and feature normalization, (ii) weight normalization, and (iii) modeling of distributional returns with an entropy-regularized MORL algorithm. The empirical results across standard continuous control benchmarks demonstrate that these changes substantially improve the quality of the produced solution sets without requiring major changes to the underlying algorithm.
Comments21 pages, 10 figures; Accepted to RLC 2026