发表机构
Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究测试监控器能否识别VLA指令中的隐含危害,发现文本守卫漏检高达95%的隐含危害请求,且激活监控和线性探针在机器人训练后区分能力下降。
AI 中文摘要
视觉-语言-动作模型(VLAs)根据指令行动而无法拒绝,因此对有害请求的筛查依赖于监控器。我们测试这些监控器是否能捕捉到因有害原因而被请求的普通机器人任务,在保持任务不变的同时,仅改变意图陈述的明确程度。$\pi_{0.5}$ 在每一级明确度下都完成任务,其频率与无害对照组相同。文本守卫几乎标记了所有直白的请求,但很少标记隐含的请求:高达95%的隐含危害运行以任务完成且未触发任何标记而结束,即使在针对机器人指令重新校准后,这一比例仍高达90%。监控模型的激活状态并不能缩小这一差距。线性探针在基础语言模型和视觉-语言模型中几乎完美地区分有害与无害指令,但在两个模型家族经过机器人训练后,这种区分能力减弱,对隐含危害尤为明显。
英文摘要
Vision-language-action models (VLAs) act on instructions without being able to refuse, so screening harmful requests falls to monitors. We test whether these monitors catch ordinary robot tasks requested for harmful reasons, holding the task fixed while varying only how explicitly the intent is stated. $π_{0.5}$ completes the task at every level of explicitness, as often as for harmless controls. Text guards flag nearly every blunt request but few implied ones: up to 95% of implied-harm runs end with the task done and no flag raised, and up to 90% even after recalibrating on robot instructions. Monitoring the model's activations does not close this gap. Linear probes separate harmful from harmless instructions almost perfectly in the base language and vision-language models, but this separation weakens after robot training in two model families, most for implied harm.
CommentsUnder review at the SPAIS 2026 workshop (CoRL 2026)