黑盒多轮交互中轨迹级安全风险预测
Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions
浏览论文内容
中文总结 AI 辅助
本文提出 Recast 框架,通过双尺度轨迹视图建模风险演变,可提前2.41轮预测88.3%的安全故障,误报率12.3%,实现轨迹级安全风险预测。
中文摘要 AI 辅助
随着大型语言模型(LLMs)从独立助手演变为自主智能体,确保其安全性需要从逐点风险评估转向理解风险如何在长时轨迹中产生与演变。在多轮交互中,恶意意图可分解至看似无害的多轮中,并通过交互轨迹逐步重构,最终引发安全故障。现有安全措施多为被动式,仅检测已显现的违规行为,缺乏预测潜在风险演变及实现 preemptive 预防的能力。为解决此局限,本文提出 Recast,一种安全风险预测框架,将 LLM 安全防护从轮级违规检测推进至轨迹级风险预测。Recast 首先通过双尺度轨迹视图从短期对话进展与长期历史上下文检索风险相关证据;接着通过捕捉当前风险配置及其时间动态,对组合式风险演变进行建模;最后,因果时间编码器学习潜在风险演变模式并预测未来风险出现轮次的分布。针对7类风险开展的大量实验表明,Recast 可预测88.3%的未来安全故障,平均提前时间为2.41轮,同时保持12.3%的误报率,彰显了轨迹级预测在安全违规发生前识别新兴风险的有效性。
英文摘要
As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk assessment to understand how risks emerge and unfold over long-horizon trajectories. In multi-turn interactions, malicious intent can be decomposed across seemingly harmless turns and gradually reconstructed through interaction trajectories, eventually resulting in safety failures. Existing safeguards remain largely reactive, detecting manifested violations while lacking the ability to predict latent risk evolution and enable preemptive prevention. To address this limitation, we propose Recast, a safety risk forecasting framework that advances LLM safeguarding beyond turn-level violation detection to trajectory-level risk prediction. Recast first retrieves risk-relevant evidence from both short-term dialogue progression and long-term historical context via a dual-scale trajectory view. It then models compositional risk evolution by capturing the current risk configuration and its temporal dynamics. Finally, a causal temporal encoder learns latent risk evolution patterns and predicts the distribution of future risk emergence turns. Extensive experiments across 7 risk categories show that Recast predicts 88.3% of future safety failures with an average lead time of 2.41 turns, while maintaining a false alarm rate of 12.3%, showcasing the effectiveness of trajectory-level forecasting in identifying emerging risks before safety violations occur.