arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越不安全检测:多轮LLM安全失败的反事实锚定证据归因

Beyond Unsafe Detection: Counterfactually Anchored Evidence Attribution for Multi-Turn LLM Safety Failures

Srinivasan Subramanian, Kazi Aminul Islam, Md. Abdullah Al Hafiz Khan

arXiv 2609.27773首次发表:更新:

发表机构

Kennesaw State University(肯尼索州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多轮LLM安全失败,提出反事实锚定的分层归因模型,识别导致不安全的轮次和标记,检测F1达0.988,且显著降低假阳性率。

AI 中文摘要

随着大语言模型(LLMs)从对话助手演进为高级智能体系统,护栏故障可能将对抗性意图转化为有害执行。然而,大多数护栏评估框架仅关注结果,评估用户请求是安全还是不安全。这种方法对于多轮失败而言是不充分的,因为对抗性意图分布在多个轮次中。这促使我们超越检测,识别将对话推向不安全轨迹的轮次和标记。为支持这一点,我们构建了一个具有行为验证和分层证据监督的多轮数据集。该数据集包含1,762个对话,包括对抗性对话、良性孪生对话和具有高风险词汇的良性变体。我们训练了一个轻量级分层归因模型,该模型预测安全违规并将其归因于贡献的用户轮次和标记跨度。该模型实现了强大的检测性能(F1=0.988),移除前15%的归因标记将对抗性分类置信度降低了51.1%。该模型在具有高风险词汇的良性对话上保持了低假阳性率,在边缘良性对话和良性高风险词汇对话上的假阳性率均低于1%,而基于关键词的表面风险基线的假阳性率分别为37.3%和94.7%。独立的人工标注支持该模型的归因性能,在84.5%的对抗性案例中,前五个归因轮次包含一个人类识别的证据承载轮次。

英文摘要

As Large Language Models (LLMs) move from conversational assistants to advanced agentic systems, guardrail failures can convert adversarial intents into harmful executions. However, most guardrail evaluation frameworks focus only on the result and assess whether a user request is safe or unsafe. This approach is insufficient for multi-turn failures, where adversarial intent is distributed across multiple turns. This motivates us to go beyond detection to identify the turns and tokens that push the conversation toward unsafe trajectories. To support this, we construct a multi-turn dataset with behavioral validation and tiered evidence supervision. The dataset contains 1,762 conversations, including adversarial conversations, benign twins, and benign variants with high-risk vocabulary. We train a lightweight hierarchical attribution model that predicts safety violations and attributes them to contributing user turns and token spans. The model achieves strong detection performance (F1=0.988), and removing the top 15% of attributed tokens reduces the adversarial classification confidence by 51.1%. The model preserves low false positive rates on benign conversations with high-risk vocabulary, with false positives below 1% on both borderline benign and benign high-risk vocabulary conversations, compared to 37.3% and 94.7% for a keyword-based surface-risk baseline. Independent human annotation supports the model's attribution performance, with the top-five attributed turns containing a human-identified evidence-bearing turn in 84.5% of adversarial cases.

CommentsThis work is currently under review for EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑