ProbGuard:基于大语言模型输出分布的校准安全风险估计
ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions
浏览论文内容
中文总结 AI 辅助
研究针对现有LLM安全防护的确定性范式局限,提出概率架构无关防护框架ProbGuard,利用早期输出分布信号校准安全风险,实现提前终止不安全输出,在9种组合设置及6种越狱攻击中均表现优异。
中文摘要 AI 辅助
近期大语言模型(LLM)安全研究广泛采用防护措施识别不安全的LLM输出。现有防护措施通常将安全评估建模为确定性分类任务,将离散的token序列映射为离散的安全标签。但该范式存在两个局限:其一,安全评估本质上是不确定问题,尤其在生成早期阶段;其二,仅依赖离散token序列会丢弃LLM输出分布中蕴含的丰富概率信息。为解决这些局限,我们提出首个完全与架构无关的概率防护框架\textsc{ProbGuard},利用LLM早期输出分布信号估计并校准安全概率,从而实现对不安全正在生成的输出的提前终止。具体而言,给定LLM生成的前缀分布,我们将安全风险建模为其后续生成动态的不安全概率,并通过蒙特卡洛采样估计该风险。通过对分布信号和校准后的安全风险进行后训练,\textsc{ProbGuard}在全部9种模型-数据集组合设置中实现了最佳校准性能,较最优基线分别降低了79.6%的平均Brier分数和71.9%的ECE。在仅观察LLM前10个解码步骤的早期输出分布后,\textsc{ProbGuard}还将6种代表性越狱攻击的攻击成功率限制在最高1%以内。
英文摘要
Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formulate safety assessment as a deterministic classification task, mapping a discrete token sequence to a discrete safety label. However, this paradigm has two limitations: First, safety assessment is inherently an uncertain problem, particularly during the early generation state. Second, relying solely on discrete token sequences discards the rich probabilistic information embedded in the LLM output distribution. To address these limitations, we propose the first completely probabilistic architecture-agnostic guardrail \textsc{ProbGuard} to leverage the LLM early output distributional signals for estimating and calibrating the safety probability, thereby enabling early stopping of unsafe ongoing outputs. Specifically, given an LLM's generated prefix distribution, we formulate the safety risk as the unsafe probability of its continued generation dynamics and estimate this risk by Monte-Carlo sampling. Through post-training on the distributional signals and calibrated safety risk, \textsc{ProbGuard} achieves the best calibration performance across all nine model--dataset combination settings, reducing the average Brier score and ECE by 79.6\% and 71.9\%, respectively, over the best baseline. \textsc{ProbGuard} further limits the attack success rate to at most 1\% across six representative jailbreak attacks after observing the LLM early output distributions from only the first ten decoding steps.