发表机构
Politecnico di Milano(米兰理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究在安全Actor-Critic最优控制中提出融合自适应经验回放与在线不确定性估计的集成架构,在二维机器人导航任务的测试中,该集成配置在极端压力测试下实现零接触且全部种子到达目标,性能优于仅移除不确定性估计的配置。
AI 中文摘要
安全Actor-Critic控制通常将障碍过滤、不确定性估计和经验回放视为独立模块,尽管每个模块都会改变学习和控制所用的数据。我们开发了一种集成架构,其中不确定性估计会更新控制障碍函数所用的障碍几何结构,过滤干预和估计残差决定回放优先级,且评论者(critic)从实际执行的动作而非标称动作中学习。我们在带有损坏障碍测量值的二维机器人导航任务中实例化该架构,并在共同的训练预算、随机种子、传感器流、探索和干扰下比较6种组件匹配的配置。评估包括中等的训练后测试、11级感知噪声扫描,以及乘数为6.0的探索性极端压力测试。在极端测试中,集成配置未记录任何接触,且在全部5个评估种子中均到达目标;其平均代价为7.63±0.44,障碍置信度的均方根误差为3.52±0.55厘米。不确定性估计的 ablation( ablation 指消融实验,即移除模型某部分以评估其作用)配置也未记录任何接触,但在5个种子中的4个到达目标,平均代价为8.96±2.08,置信度误差为11.08±1.23厘米。有限训练边界明确了回放暴露,而鲁棒障碍条件则阐明了所需的估计误差和可行性假设。这些结果支持在该基准上耦合估计、安全过滤和回放;更广泛的安全性和收敛性主张有待进一步研究。
英文摘要
Safe actor-critic control often treats barrier filtering, uncertainty estimation, and experience replay as separate modules, even though each changes the data used for learning and control. We develop an integrated architecture in which the uncertainty estimate updates the obstacle geometry used by a control barrier function, filter interventions and estimation residuals determine replay priority, and the critic learns from the executed rather than nominal action. We instantiate the architecture on a two-dimensional robot-navigation task with corrupted obstacle measurements and compare six component-matched configurations under common training budgets, random seeds, sensor streams, exploration, and disturbances. Evaluation includes a moderate post-training test, an eleven-level perception-noise sweep, and an exploratory extreme-stress test at multiplier $6.0$. In the extreme test, the integrated configuration recorded no contacts and reached the goal in all five evaluation seeds. Its mean cost was $7.63\pm0.44$ and its obstacle-belief root-mean-square error was $3.52\pm0.55$ cm. The uncertainty-estimation ablation also recorded no contacts but reached the goal in four of five seeds, with mean cost $8.96\pm2.08$ and belief error $11.08\pm1.23$ cm. A finite-training bound clarifies replay exposure, and a robust barrier condition states the required estimation-error and feasibility assumptions. The results support coupling estimation, safety filtering, and replay on this benchmark; broader safety and convergence claims require further study.
CommentsCode, deterministic seeds, data, figures, and protocol files are archived at https://doi.org/10.5281/zenodo.21515850 and https://github.com/SDNT8810/safe-actor-critic-aer-ue-reproducibility