SingProbe 技术报告
SingProbe Technical Report
浏览论文内容
中文总结 AI 辅助
SingProbe 是一种轻量级 LLM 内在运行时护栏,复用 LLM 隐藏状态以 token 级别预测意图、安全性与幻觉风险,仅约 200 万参数、额外开销<0.5%,还可指导安全解码,其衍生的 SingProbe-Med 可用于医疗生成的风险干预。
中文摘要 AI 辅助
运行时护栏对于可靠部署大型语言模型(LLM)至关重要,但现有方法通常依赖独立的外部模型,这会带来额外的推理成本、延迟的安全信号,以及与能力日益增强的基础模型之间的能力不匹配。为解决这些问题,我们提出了 SingProbe,这是一种轻量级的内在运行时护栏,可直接复用 LLM 推理过程中产生的隐藏状态,并与自回归解码协同运行。在统一框架内,SingProbe 以 token 级别持续预测查询意图、响应安全性和幻觉风险,额外的护栏推理开销可忽略不计,提供了一种“免费午餐”式的解决方案。我们还推出了 SingStreamBench,这是一个基准,用于评估流式护栏在良性前缀上保持不激活,同时及时检测新出现的不安全内容的能力。大量实验表明,SingProbe 仅具有约 200 万参数,额外开销低于 0.5%,与规模大得多的独立护栏和专门的幻觉检测器相比,实现了相当或更优的性能。除了被动检测,我们还表明 SingProbe 的评分可以预测未来生成风险并指导受限的安全解码。我们进一步将该范式扩展到医疗生成,推出了 SingProbe-Med,它仅在出现临床相关风险时才选择性激活风险导向的解码干预。这些结果共同表明,内部模型表示为生成时的监控和控制提供了有效且高效的接口。
英文摘要
We present SingProbe, an open intrinsic guardrail framework for generation-time monitoring of LLMs. Intrinsic guardrails reuse hidden states already produced by the base model during autoregressive decoding, rather than relying on an independent model to repeatedly process generated text. While this route has been explored in industrial systems, the community lacks a broadly reusable open stack that combines cross-model guard adaptations, unified training methods, serving integrations, and systematic evaluation resources. SingProbe is designed to provide this missing layer and uses a lightweight probe to continuously produce query-intent, response-safety, and hallucination-risk signals during decoding. This report describes the full intrinsic-guardrail stack: training methods, serving integrations with SGLang and vLLM, and adapted guard models for 29 open-source base models across diverse families and scales. We also introduce SingStreamBench, a benchmark that measures whether streaming guardrails remain inactive on benign prefixes while promptly detecting emerging unsafe content. Across evaluations of safety, streaming detection, hallucination detection, false-positive robustness, online monitoring, and runtime overhead, SingProbe provides performance competitive with, and in several settings stronger than, state-of-the-art standalone guardrails and specialized hallucination detectors, while adding less than 0.5% serving overhead in our implementation. Beyond passive monitoring, we show that intrinsic guard signals can guide constrained decoding and selectively activate medical-risk interventions in SingProbe-Med. By open-sourcing our infrastructure, training methods, and model adaptations, we aim to facilitate the broader adoption and deployment of intrinsic guardrails, as well as further research in this direction.
发表机构
- Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。