发表机构
College of Medicine and Public Health, Flinders University(弗林德斯大学医学与公共卫生学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多轮大语言模型系统的对话风险累积问题,提出会话层CRA框架,跟踪三个轨迹信号,提供评分方法,发布多个基准测试集及评估协议,重点在于分布内会话评分,为大语言模型安全防护提供新方法。
AI 中文摘要
大多数大语言模型(LLM)的安全防护栏孤立地评估每个提示-响应对,这忽略了仅在对话中出现的故障,因为良性轮次可能累积成危害。我们将此称为对话风险累积(CRA),包括逐渐的意图漂移、禁止指令的碎片化组合以及重复披露导致的敏感度增加。我们提出了一个会话层CRA框架,它跟踪三个轨迹信号:来自会话锚点的语义漂移、提取实体上的敏感度加权信息累积图以及捕获合规意愿增加的合规梯度信号。在评分方面,我们提供了用于归因和消融的无监督凸融合,以及CRA-Net DA,这是一个紧凑的学习轨迹模型,通过家族对抗目标进行训练以减少长度和主题覆盖混淆。为了对CRA进行基准测试,我们发布了CRA-Bench v0.1(跨越三个威胁家族的1200个八轮会话以及主题匹配的良性双胞胎)、CRA-Bench v0.2(LLM释义变体以减少模板伪影)以及一个扩展的五家族集(2000个会话,增加了角色引导和上下文填充)。我们引入了一种轨迹原生评估协议,包括会话级分割、混合集阈值校准、轨迹AUROC、检测轮数、校准的误报指标、引导置信区间、留一法家族外诊断压力测试以及合成到人类的转移检查。研究重点在于CRA-Bench和人类转移子集上的分布内会话评分。
英文摘要
Most safety guardrails for large language models (LLMs) evaluate each prompt-response pair in isolation, which misses failures that arise only over a dialogue as benign turns compose into harm. We term this Conversational Risk Accumulation (CRA): gradual intent drift, fragmented assembly of prohibited instructions, and sensitivity build-up from repeated disclosures. We propose a session-layer CRA Framework that tracks three trajectory signals: semantic drift from a session anchor, a sensitivity-weighted information accumulation graph over extracted entities, and a compliance-gradient signal capturing increasing willingness to comply. For scoring, we provide (i) an unsupervised convex fusion for attribution and ablations, and (ii) CRA-Net DA, a compact learned trajectory model trained with family-adversarial objectives to reduce length and topic-coverage confounds. To benchmark CRA, we release CRA-Bench v0.1 (1,200 eight-turn sessions across three threat families with topic-matched benign twins), CRA-Bench v0.2 (LLM-paraphrased variants to reduce template artifacts), and an extended 5-family set (2,000 sessions adding persona priming and context stuffing). We introduce a trajectory-native evaluation protocol with session-level splits, mixed-set threshold calibration, Trajectory AUROC, turns-to-detection, calibrated false-positive metrics, bootstrap confidence intervals, leave-one-family-out diagnostic stress tests, and synthetic-to-human transfer checks. Claims focus on within-distribution session scoring on CRA-Bench and human-transfer subsets.
Comments45 pages, 10 figures, 20 tables