arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19472cs.AIcs.CLcs.CRcs.LG

超越界面的安全性:通过大语言模型的潜在状态检测危害

Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models

Alizishaan Khatri, Chiquita Prabhu, Omkar Neogi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过从LLaMA-3.1-8B提取激活并训练轻量级MLP探针,实现了与大型护栏模型相当的有害提示检测性能,同时显著降低延迟和计算开销。

中文摘要 AI 辅助

自主系统日益依赖大语言模型(LLMs),然而围绕这些模型的安全基础设施引入了延迟和计算开销。这限制了在资源受限、时间关键型部署中的实用性。现有的外部护栏模型对模型的内部运作视而不见,造成了根本性的保障缺口。我们提出疑问:模型是否已经知道内容何时有害?我们从LLaMA-3.1-8B中提取激活,并训练轻量级MLP分类器探针(1260万参数)来检测有害提示。在WildJailbreak、Beavertails和AEGIS 2.0上评估,我们的探针分别达到99%、83%和84%的F1分数,与体积大1000倍的护栏模型相比具有竞争力,同时降低了延迟和计算成本。

英文摘要

Autonomous systems increasingly rely on Large Language Models (LLMs) yet the safety infrastructure surrounding these models introduces latency and compute overhead. This limits utility in resource-constrained, time-critical deployments. Existing external guardrail models remain blind to the model's internal workings, creating a fundamental assurance gap. We ask: does the model already know when the content is harmful? We extract activations from LLaMA-3.1-8B and train lightweight MLP classifier probes (12.6M parameters) to detect harmful prompts. Evaluated on WildJailbreak, Beavertails, and AEGIS 2.0, our probes achieve F1 scores of 99%, 83%, and 84%, respectively competitive with 1000x larger guard models while cutting latency and compute costs.

发表机构

  • Wrynx Inc.(Wrynx公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑