发表机构
King Abdullah University of Science and Technology(阿卜杜拉国王科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出基于证据的安全评估基准TRACE,覆盖大型推理模型的完整推理流程,评估发现现有防护栏模型对推理轨迹的安全判断难度大,且难以准确提取支撑证据,凸显了改进防护栏模型的必要性。
AI 中文摘要
大型推理模型(Large Reasoning Models, LRMs)会生成中间推理轨迹,即便其最终响应看似安全,轨迹中也可能包含不安全内容。防护栏(Guardrail)模型旨在检测并拦截不安全内容,但现有的不安全内容检测基准主要聚焦于提示词和最终响应,基本未对推理轨迹进行考察;此外,这些基准通常仅提供二元安全标签,未提供支撑判断的证据标注。为解决这些局限,我们推出了TRACE——一个基于证据的安全评估基准,覆盖LRM的完整推理流程:提示词、推理轨迹和最终响应。TRACE包含两种语言的提示词,覆盖9类风险和10种攻击策略;针对每个提示词,4个LRMs会生成推理轨迹和最终响应,我们对每个组件的安全性进行标注,并从对应源文本中提取支撑证据。在TRACE上评估18个防护栏模型的结果显示,对推理轨迹的安全判断比对提示词或最终响应的判断难度大得多,且当前模型难以准确提取支撑证据。这些发现凸显了对防护栏模型的需求,这类模型需能在LRM的完整推理流程中可靠检测并精确定位不安全内容。
英文摘要
Large Reasoning Models (LRMs) generate intermediate reasoning traces that may contain unsafe content, even when their final responses appear safe. Guardrail models are designed to detect and block unsafe content, yet existing benchmarks for unsafe content detection focus primarily on prompts and final responses, leaving reasoning traces largely unexamined. Moreover, these benchmarks typically provide only binary safety labels, without evidence annotations that justify the judgments. To address these limitations, we introduce TRACE, an evidence-grounded safety evaluation benchmark that covers the entire LRM inference pipeline: prompts, reasoning traces, and final responses. TRACE includes prompts in two languages spanning nine risk categories and ten attack strategies. For each prompt, four LRMs generate reasoning traces and final responses, and we annotate the safety of each component and extract supporting evidence from the corresponding source text. Evaluating 18 guardrail models on TRACE reveals that safety judgment for reasoning traces is substantially more challenging than for prompts or final responses, and that current models struggle to accurately extract supporting evidence. These findings highlight the need for guardrail models that can reliably detect and precisely localize unsafe content across the LRM inference pipeline.
CommentsEMNLP 2026 Main