AI 中文总结
研究非完整约束下无模型强化学习奖励函数设计难题,提出含覆盖门控对齐反馈等的参数化奖励塑造框架,通过联合元优化及基于代理的贝叶斯优化,优化后的深度Q网络代理在停车任务中表现出色。
AI 中文摘要
为非完整约束下的无模型强化学习设计有效奖励函数仍是一个持续挑战,常导致严重局部极小值。本文提出参数化奖励塑造框架,含覆盖门控对齐反馈、驱动方向切换正则化等,并在自主平行停车任务中评估。关键是表明环境奖励参数和算法超参数深度相互依赖,需联合元优化。基于代理的贝叶斯优化的深度Q网络代理解决了控制失败模式,在成功率和轨迹平滑度上显著优于未校准基线。
英文摘要
Designing effective reward functions for model-free reinforcement learning under non-holonomic constraints remains a persistent challenge, often resulting in severe local minima such as policy paralysis or over-conservative hazard avoidance. In this work, we present a parameterized reward shaping framework featuring coverage-gated alignment feedback, drive-direction switch regularization, and an aligned episode termination mechanism evaluated on an autonomous parallel parking task. Crucially, we show that environmental reward parameters and algorithmic hyperparameters are deeply co-dependent, requiring joint meta-optimization to achieve stable convergence. By employing surrogate-based Bayesian optimization, our co-optimized Deep Q-Network (DQN) agent resolves characteristic control failure modes, significantly outperforming uncalibrated baselines across both success rate and trajectory smoothness.
Comments12 pages, 8 figures. Includes supplementary video demonstration and open-source code link