AI 中文总结
该研究提出受Linux安全模块启发的LMSM框架,通过分离调解与策略实现LLM安全,在Qwen3-4B上可降低攻击成功率,同时保持较高吞吐量。
AI 中文摘要
大语言模型(LLM)越来越多地采用分层防御部署,但恶意提示仍可绕过这些防御。可解释性方法可揭示模型在生成路径上的内部信号,这些信号可用于实施防御,但这些信号本身并非安全控制措施。为安全目的对其进行适配的部署通常会将每个信号与其自身的校准、策略逻辑和干预代码耦合,因此每个新的人工制品都会产生集成工作,而非强化共享防御。我们提出Language Model Security Modules(LMSM,语言模型安全模块),这是一种将Linux安全模块(LSM)背后的分离机制适配到LLM服务的安全框架。在LMSM中,选定的安全后端会暴露校准后的证据,带版本的策略会在可信的每请求上下文上评估活跃规则,而单独的网关会授权缓冲输出的释放。该设计将调解正确性与策略有效性分离,允许在不重建请求处理或实施机制的情况下更改后端、规则或调度。我们的原型表明该分离机制在实践中有效:使用Hugging Face Transformers和连续批处理的vLLM,同一底层平台可支持基于人工制品的稀疏自编码器(SAE)、转码器部署以及任务适配的密集探针,在调度器波动下保留特定请求的决策,并对每个请求选择性实施和组合多个规则。在Qwen3-4B上,LMSM-Checkpoint将HarmBench攻击成功率从39.20%降至3.32%,XSTest的错误弃权(不执行)率从2.40%升至4.40%,同时在32个活跃序列下保留了无监控工作的匹配服务路径98.14%的吞吐量。LMSM为可解释性和模型内部分析的进展提供了运行时实施的通用路径。
英文摘要
Large language models (LLMs) are increasingly deployed with layered defenses, yet malicious prompts can still bypass them. Interpretability methods can expose model-internal signals along the generation path that could inform enforcement, but these signals are not security controls by themselves. Deployments that adapt them for safety typically couple each signal to its own calibration, policy logic, and intervention code, so each new artifact creates integration work instead of strengthening a shared defense. We present Language Model Security Modules (LMSM), a security framework that adapts the separation behind Linux Security Modules (LSM) to LLM serving. In LMSM, a selected security backend exposes calibrated evidence, a versioned policy evaluates active rules over trusted per-request context, and a separate gate authorizes buffered output release. This design separates mediation correctness from policy effectiveness, and it allows backend, rule, or schedule changes without rebuilding request handling or enforcement. Our prototype shows the separation working in practice: with Hugging Face Transformers and continuously batched vLLM, the same substrate hosts artifact-backed sparse autoencoder (SAE) and transcoder deployments and task-fitted dense probes, preserves request-specific decisions under scheduler churn, and selectively enforces and composes multiple rules per request. On Qwen3-4B, LMSM-Checkpoint reduces HarmBench attack success rate from 39.20% to 3.32%, with XSTest false refusals rising from 2.40% to 4.40%, while retaining 98.14% of the throughput of a matched serving path that performs no monitoring work at 32 active sequences. LMSM gives advances in interpretability and model-internal analysis a common path to runtime enforcement.