arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15218cs.AIcs.CR

当言语安全但行动致命:在隐藏状态风险空间中探究超越文本安全的物理危险

When Words Are Safe But Actions Kill: Probing Physical Jailbreak Beyond Textual Jailbreak in Hidden-State Risk Space

Weimeng Wang, Ziqiang Wang, Zihang Zhan, Chuanpu Fu, Qi Li, Ke Xu

首次发表
浏览论文内容

中文总结 AI 辅助

研究大语言模型中语言无害指令在现实世界中的物理危险与文本级内容危险是否相同,提出PRISM方法,通过隐藏状态方向分析等证明CD和PD可分离,PRISM在多个基准测试中表现良好,能检测物理危险而非仅依赖明确不安全措辞。

中文摘要 AI 辅助

大语言模型越来越多地充当具身智能体的高级规划器,语言上无害的指令一旦在现实世界中落地就可能变得不安全。研究这种现实世界中的危险是否与普通文本级内容危险是同一安全问题。通过隐藏状态方向分析和随机分割零测试表明,内容危险(CD)和物理危险(PD)在Qwen2.5 - 3B/7B/14B/32B、Phi - 3.5和SmolLM2等语言模型表示中形成可分离信号。在此基础上提出PRISM,在SafeAgentBench上准确率达86.2 - 87.7%,误报率11.7 - 13.7%,还引入PhysicalSafetyBench - 1K进行测试,PRISM在该测试中准确率达99.6%,误报率0.7%,且在SafeText和EARBench上也有良好表现。

英文摘要

Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world. We study whether this physically grounded jailbreak is the same safety problem as ordinary textual jailbreak. Through hidden-state direction analysis and random-split null tests, we show that textual jailbreak (TJ) and physical jailbreak (PJ) form separable signals in LLM representations across Qwen2.5-3B/7B/14B/32B, Phi-3.5 and SmolLM2. Building on this separability, we propose PRISM, a single-layer L2-regularized logistic probe over full hidden states. PRISM achieves 86.2--87.7\% accuracy on SafeAgentBench with 11.7--13.7\% false-positive rates (FPRs), while same-scale LLM judges over-block safe tasks at 24.7--39.0\% FPR. To test whether the result survives lexical-shortcut controls, we introduce an interaction-balanced revision of PhysicalJailbreakBench-2K (PJB-2K): a fixed 2{,}000-row comparison set sampled by label and physical mechanism from a larger object--site construction. On the underlying 10{,}000-row pool, word-TFIDF and the embedding layer remain at chance (AUC 0.497 and 0.500). At layer 25, selected by an i.i.d. sweep, cell-grouped cross-validation gives PRISM 0.718 AUC, compared with 0.398 for a physics-free label control under the same protocol. On the identical 2{,}000 comparison rows, these PRISM predictions obtain 0.671 balanced accuracy, while Qwen2.5 judges from 3B to 72B obtain 0.538--0.577 and exhibit high FPR. These results support hidden-state probing as a representation-level method for physical safety beyond text moderation, without relying on the near-perfect scores of shortcut-prone paired templates.

发表机构

  • Tsinghua University(清华大学)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑