解构演员-评论家算法:面向从业者的设计组件大规模实证研究
Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners
浏览论文内容
中文总结 AI 辅助
研究针对演员-评论家算法在现实世界控制场景中的应用,通过实际水处理厂控制任务进行超33000次实验,分析算法设计组件对运行变异性和超参数敏感性的影响,为从业者在新场景中进行组件级决策提供实证指导。
中文摘要 AI 辅助
强化学习越来越多地被用于控制现实世界系统,如核聚变等离子体、自动驾驶车辆、药物发现和饮用水处理等,这些场景中可靠性至关重要且调整预算有限。演员-评论家算法有一系列设计决策,如策略更新方式等。通过源自实际水处理厂的控制任务,我们分析超33000次实验,以确定这些组件如何影响运行中的变异性和对超参数的敏感性。常见默认设置可靠性低,而有界分布和自适应更新策略在多种设置下保持稳健。这些发现为从业者提供了实证指导。
英文摘要
Reinforcement learning is increasingly being considered for controlling real-world systems, from fusion plasma and autonomous vehicles to drug discovery and drinking water treatment, where reliability is essential and tuning budgets are limited. Actor-critic algorithms share a set of design decisions, such as how the policy is updated, how it represents the distribution over actions, how its gradient is estimated, and how often it is updated relative to the value estimator. Using a control task derived from a real water treatment plant, we analyze over 33,000 experiments to determine how these components affect variability across runs and sensitivity to hyperparameters. Common defaults, such as Gaussian action distributions with pathwise gradient estimators, are among the least reliable configurations, whereas bounded distributions with adaptive update schedules remain robust across a wide range of settings. These findings offer empirical guidance to practitioners across scientific and engineering domains for understanding and making component-level decisions when adapting actor-critic methods to new real-world control settings.