大型语言模型 (LLMs) 越来越多地充当具体代理的高级规划器,,其中语言上良性的指令一旦扎根于物理世界就可能变得不安全。我们研究这种基于物理的越狱是否与普通的文本越狱存在相同的安全问题。通过隐藏状态方向分析和随机分割空测试,,我们表明文本越狱(TJ)和物理越狱(PJ)在Qwen2.5-3B/7B/14B/32B, Phi-3.5和SmolLM2的LLM表示中形成可分离的信号。基于这种可分离性,,我们提出 PRISM, 是一种针对完全隐藏状态的单层 L2 正则化逻辑探测。 PRISM 在 SafeAgentBench 上达到 86.2--87.7\% 的准确率,误报率为 11.7--13.7\% (FPRs),,而同规模的 LLM 判断超块安全任务的 FPR 为 24.7--39.0\% FPR。为了测试结果是否能经受住词汇快捷方式控件,的考验,我们引入了PhysicalJailbreakBench-2K (PJB-2K):的交互平衡修订版,这是一个固定的2{,}000行比较集,通过来自更大对象(站点构建)的标签和物理机制进行采样。在底层 10{,}000 行池, 字 TFIDF 和嵌入层上,概率仍为 (AUC 0.497 和 0.500)。在由 i.i.d. 选择的 25, 层scan, 细胞分组交叉验证得出 PRISM 0.718 AUC,,而相同方案下无物理标签对照的 AUC, 为 0.398。在相同的 2{,}000 比较行, 上,这些 PRISM 预测获得 0.671 平衡精度,,而 Qwen2.5 从 3B 到 72B 的判断获得 0.538--0.577 并表现出高 FPR。这些结果支持隐藏状态探测作为超越文本审核, 的物理安全的表示级方法,而不依赖于容易出现快捷方式的配对模板的近乎完美的分数。
Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world. We study whether this physically grounded jailbreak is the same safety problem as ordinary textual jailbreak. Through hidden-state direction analysis and random-split null tests, we show that textual jailbreak (TJ) and physical jailbreak (PJ) form separable signals in LLM representations across Qwen2.5-3B/7B/14B/32B, Phi-3.5 and SmolLM2. Building on this separability, we propose PRISM, a single-layer L2-regularized logistic probe over full hidden states. PRISM achieves 86.2--87.7\% accuracy on SafeAgentBench with 11.7--13.7\% false-positive rates (FPRs), while same-scale LLM judges over-block safe tasks at 24.7--39.0\% FPR. To test whether the result survives lexical-shortcut controls, we introduce an interaction-balanced revision of PhysicalJailbreakBench-2K (PJB-2K): a fixed 2{,}000-row comparison set sampled by label and physical mechanism from a larger object--site construction. On the underlying 10{,}000-row pool, word-TFIDF and the embedding layer remain at chance (AUC 0.497 and 0.500). At layer 25, selected by an i.i.d. sweep, cell-grouped cross-validation gives PRISM 0.718 AUC, compared with 0.398 for a physics-free label control under the same protocol. On the identical 2{,}000 comparison rows, these PRISM predictions obtain 0.671 balanced accuracy, while Qwen2.5 judges from 3B to 72B obtain 0.538--0.577 and exhibit high FPR. These results support hidden-state probing as a representation-level method for physical safety beyond text moderation, without relying on the near-perfect scores of shortcut-prone paired templates.
科目: 人工智能 (cs.AI); 密码学和安全 (cs.CR)
Subjects: Artificial Intelligence (cs.AI); Cryptography and Security (cs.CR)