约束强化学习中的视觉-语言信号:无需预判的安全增益
Vision--Language Signals in Constrained RL: Safety Gains Without Anticipation
浏览论文内容
中文总结 AI 辅助
本文提出VLM-Safe-RL框架,将冻结的CLIP信号集成到PPO-Lagrangian中,在MetaDrive Hard上使灾难率从31.6%降至19.4%,但分析显示该信号不预判碰撞,安全提升源于其他机制。
中文摘要 AI 辅助
安全强化学习旨在最大化任务性能的同时满足安全约束。然而,在驾驶基准测试中,碰撞成本通常仅在碰撞发生时出现,无法对即将到来的危险提供预警。冻结的视觉-语言模型可以提供密集的语义反馈,但其分数是否能预判碰撞,以及哪个组件驱动了观察到的安全改进,仍不清楚。此外,回合制成本也可能偏向于任务进展甚微的策略。为解决这些问题,我们提出了VLM-Safe-RL框架,该框架通过奖励塑形和增强的乘子更新,将冻结的CLIP信号集成到PPO-Lagrangian中。在MetaDrive Hard(结合了最密集交通和最大地图)上,灾难率从31.6%降至19.4%。FormulaOne-L2分析发现,没有证据表明CLIP信号能预判碰撞,且VLM项对拉格朗日乘子的影响可忽略不计。这些发现表明,观察到的灾难率降低是有条件的,且没有碰撞预判的证据。
英文摘要
Safe reinforcement learning seeks policies that maximise task performance while satisfying safety constraints. In driving benchmarks, however, collision costs typically appear only at the time of collision, providing no advance warning of an approaching hazard. Frozen vision--language models can provide dense semantic feedback, yet it remains unclear whether their scores anticipate collisions and which component drives an observed safety improvement. Episodic cost can also favour policies that make little task progress. To address these gaps, we propose VLM-Safe-RL, a framework that integrates frozen CLIP signals into PPO-Lagrangian through reward shaping and an augmented multiplier update. On MetaDrive Hard, which combines the densest traffic with the largest map, the catastrophe rate falls from 31.6\% to 19.4\%. FormulaOne-L2 analysis finds no evidence that the CLIP signals anticipate collisions and shows that the VLM term has a negligible effect on the Lagrange multiplier. These findings show a conditional reduction in observed catastrophe rate without evidence of collision anticipation.
发表机构
- Iowa State University(爱荷华州立大学)
机构由 AI 辅助整理,请以论文原文为准。