arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于DMD的提示-响应嵌入动态分类实现大语言模型安全

Enforcing LLM Safety through DMD-based Classification of Prompt-Response Embedding Dynamics

Mohamed Akrout, Olivera Kotevska, Dan Wilson

arXiv 2608.19579首次发表:更新:

发表机构

University of Tennessee; Oak Ridge National Laboratory(田纳西大学; 橡树岭国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文将幻觉检测的动力系统框架扩展到LLM安全分类,通过拟合安全与不安全状态的Koopman模型,利用差分残差评分分类不安全输出,在多基准测试中验证了纳入提示嵌入的改进效果。

AI 中文摘要

大语言模型(LLM)越来越多地被部署在高风险应用中,但其生成有毒、有害或违反政策内容的倾向带来了重大风险。以黑盒方式高效检测这些不安全输出仍然是一个未解决的挑战。在本文中,我们将最近提出的用于幻觉检测的动力系统框架扩展到LLM安全分类中。通过将提示和响应都投影到高维嵌入空间,并为安全和不安全状态分别拟合基于Koopman的预测模型,我们使用新的差分残差评分来分类新输出,该评分比较安全和不安全状态的预测误差。一个关键贡献是纳入了提示和响应嵌入动态,产生的拟合Koopman算子捕获了关键的交互模式。我们在三个安全基准上使用三个嵌入模型评估了我们的黑盒方法。我们的结果表明,纳入提示嵌入会带来一致的改进,特别是对于与交互相关的违规,当与因果解码器(例如,在Llama-3中)配对时;而仅响应相关的违规则更受益于密集语义嵌入表示。这些发现为使用动力系统分析AI系统打开了大门,而非采用用AI对动力系统进行建模的主导范式。

英文摘要

Large Language Models (LLMs) are increasingly deployed in high-stakes applications, yet their tendency to generate toxic, harmful, or policy-violating content poses significant risks. Detecting these unsafe outputs efficiently in a black-box manner remains an open challenge. In this paper, we extend a recently proposed dynamical systems framework designed for hallucination detection to LLM safety classification. By projecting both prompts and responses into high-dimensional embedding spaces and fitting separate Koopman-based predictive models for safe and unsafe regimes, we classify new outputs using a new differential residual score that compares prediction errors of the safe and unsafe regimes. A key contribution is the incorporation of the prompt and response embedding dynamics, yielding fitted Koopman operators that capture crucial interaction patterns. We evaluate our black-box method across three safety benchmarks using three embedding models. Our results show that incorporating prompt embeddings yields consistent improvements, particularly for interaction-dependent violations when paired with causal decoders (e.g., in Llama-3), while response-only violations benefit more from dense semantic embedding representations. These findings opens the door for using dynamical systems to analyze AI systems rather than the dominant paradigm of using AI to model dynamical systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑