arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20820cs.AI

通过组合边界与安全持久性实现大语言模型安全的多轮认证鲁棒性

Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence

发表机构北京大学 · 中国人民大学 · 清华大学
另 2 家 · 查看机构详情
  • Peking University(北京大学)
  • Renmin University of China(中国人民大学)
  • Tsinghua University(清华大学)
  • University of Chinese Academy of Sciences(中国科学院大学)
  • Tencent Hunyuan(腾讯混元)

机构由 AI 辅助整理,请以论文原文为准。

Yang Liu, Bin Chong, Wenkai Yang, Shuai Zhang, Yancheng Chen, Feiyu Han, GuoZhen, Cheng Zhang, Huaibing Xie, Changze Lv, Shihan Dou, Pluto Zhou

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLMs易受多轮越狱攻击的问题,提出MTCR框架,通过状态对抗MDPs等方法实现多轮认证鲁棒性,实验表明其经验安全性能超认证边界。

中文摘要 AI 辅助

大语言模型(LLMs)易受多轮越狱攻击,这类攻击会逐步操纵对话上下文。现有认证鲁棒性方法仅适用于单轮输入,而直接的多轮组合会产生随轮数呈指数级下降的边界。本文提出多轮认证鲁棒性(MTCR)框架,该框架通过状态对抗马尔可夫决策过程(State-Adversarial MDPs)对对话安全进行建模,并将k轮认证鲁棒性定义为k次对抗轮次下的最坏情况安全概率。MTCR包含四个部分:(i)通过嵌入空间模式分解实现组合认证,产生比直接相乘更紧的认证下界;(ii)(α,β)安全持久性,将退化率从p^k提升至β^k(β>p),并给出可解释的轮数估计;(iii)匹配的信息论上界,确立边界的紧致性;(iv)结合上述结果的统一算法。在6个LLMs上针对ε有界攻击和Crescendo式攻击的实验证实,经验安全性能始终超过认证边界。

英文摘要

Large language models (LLMs) are vulnerable to multi-turn jailbreak attacks that progressively manipulate conversation context. Existing certified robustness methods are limited to single-turn inputs; naive multi-turn composition yields bounds that degrade exponentially in the number of turns. We introduce Multi-Turn Certified Robustness (MTCR), a framework that models conversational safety via State-Adversarial MDPs and defines $k$-turn certified robustness as the worst-case safety probability across $k$ adversarial turns. MTCR comprises: (i) compositional certification via embedding-space mode decomposition, yielding tighter certified lower bounds than naive multiplication; (ii) $(α,β)$-safety persistence, improving the degradation rate from $\underline{p}^{k}$ to $β^k$ (with $β> \underline{p}$) and yielding interpretable horizon estimates; (iii) matching information-theoretic upper bounds establishing tightness; and (iv) a unified algorithm combining these results. Experiments on six LLMs under $ε$-bounded and Crescendo-style attacks confirm that empirical safety consistently exceeds the certified bounds.

↑