ForeSight:通过早期安全信号蒸馏增强风险监控
ForeSight: Enhancing Risk Monitoring via Early Safety Signal Distillation
浏览论文内容
中文总结 AI 辅助
本文提出ForeSight框架,利用首令牌隐藏状态蒸馏早期安全信号,实现对大型语言模型输出风险的早期高效预测,在多个安全基准上验证了其优越性。
中文摘要 AI 辅助
随着大型语言模型(LLMs)的日益部署,有害内容的生成已成为一个关键的安全问题。现有的安全措施在输入、输出或流式生成阶段运作,而依赖表面标记或输出逻辑的早期风险方法可能因初始信号较弱而失效,使用密集表示的内部检测器可能保留高度纠缠且冗余的安全无关信息。因此,目前尚不清楚最早的后生成隐藏状态是否已经包含关于最终响应有害性的可靠信号。为解决这一空白,我们提出了ForeSight,一种首令牌输出风险预测框架,将微弱且冗余的早期安全信号蒸馏为紧凑的、层感知的风险表示。在五个安全基准和两个目标模型上的实验表明,ForeSight在仅依赖首令牌隐藏状态的情况下,实现了优越且高效的早期风险预测。代码可在以下网址获取:this https URL
英文摘要
As large language models (LLMs) are increasingly deployed, the generation of harmful content has become a critical safety concern. Existing safeguards operate at the input, output, or streaming-generation stages, while early-risk methods that rely on surface tokens or output logits may suffer from weak initial signals, and internals-based detectors using dense representations may retain highly entangled and redundant safety-irrelevant information. It therefore remains unclear whether the earliest post-generation hidden states already contain reliable signals about final-response harmfulness. To address this gap, we propose ForeSight, a first-token output-risk forecasting framework that distills weak and redundant early safety signals into compact, layer-aware risk representations. Experiments on five safety benchmarks and two target models demonstrate that ForeSight achieves superior and efficient early-risk forecasting while relying solely on first-token hidden states. The code is available at: https://github.com/Scabbards1500/Foresight
发表机构
- University of California, San Diego(加利福尼亚大学圣迭戈分校)
- East China Normal University(华东师范大学)
- Shanghai University(上海大学)
- Zhejiang University(浙江大学)
机构由 AI 辅助整理,请以论文原文为准。