arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

并非看见危险:冻结的视觉-语言安全分数衡量的是其字幕库

It Is Not Seeing the Hazard: A Frozen Vision-Language Safety Score Measures Its Caption Bank

Samuel Tetteh, Cody Fleming

arXiv 2610.09517首次发表:更新:

发表机构

Iowa State University(爱荷华州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过受控实验评估冻结CLIP安全分数,发现其下降主要反映与字幕库的场景相似度而非真实危险,策略收益不能证明危险感知能力。

AI 中文摘要

冻结的视觉-语言模型越来越多地为强化学习提供安全信号。其使用前提是,与描述危险的文本的相似度即指示危险本身。然而,策略回报和碰撞率无法揭示一个分数究竟是检测到了危险,还是对场景的相关特征做出响应。基于VLM的方法通过将图像-文本相似度转化为奖励、成本或置信权重,在驾驶和安全强化学习基准测试中报告了性能提升。这类信号有望减少对人工设计反馈的依赖。但它们也可能反映提示结构、嵌入几何或相机视角,从而使其安全含义未经证实。为弥补这一空白,我们对一个冻结的CLIP提示边际安全分数进行了受控评估。我们将该分数应用于从未接收该分数的策略所生成的轨迹,将接触前观测与具有可比危险几何的无接触观测进行匹配,并改变字幕、编码器和相机视角。在三个策略、180个回合和130次孤立的接触起始事件中,该分数在接触前约二十步时下降。机制控制表明,该分数主要追踪其字幕所共享的场景的相似度,并随字幕分离度和相机视角而变化。一个恒定置信度控制保留了较低的灾难率点估计,因此策略收益并不能确立危险感知。

英文摘要

Frozen vision-language models increasingly provide safety signals for reinforcement learning. Their use assumes that similarity to language describing danger indicates the hazard itself. Yet policy return and collision rate cannot reveal whether a score detects hazards or responds to correlated features of the scene. VLM-based methods have reported gains in driving and safe-RL benchmarks by converting image-text similarity into rewards, costs, or confidence weights. Such signals promise to reduce reliance on manually designed feedback. They may also reflect prompt structure, embedding geometry, or camera viewpoint, leaving their safety meaning unverified. To address this gap, we present a controlled evaluation of a frozen CLIP prompt-margin safety score. We apply the score to trajectories generated by policies that never receive it, match pre-contact observations to contact-free observations with comparable hazard geometry, and vary the captions, encoder, and camera view. Across three policies, 180 episodes, and 130 isolated contact onsets, the score decreases for about twenty steps before contact. Mechanism controls indicate that the score mainly tracks resemblance to the scene shared by its captions and changes with caption separation and camera view. A constant-confidence control retains the lower catastrophe-rate point estimate, so policy gains do not establish hazard perception.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑