MJ:通过分解信用分配实现多轮语言模型越狱
MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment
浏览论文内容
中文总结 AI 辅助
研究针对多轮语言模型越狱的信用分配难题,提出DC-GRPO框架,通过结合即时和未来信用为各轮次分配学习信号,经静态和动态加权规则实例化,在多模型和基准测试中显著优于现有方法,核心优势在于轮次级组相对信用分配。
中文摘要 AI 辅助
现代大语言模型在交互式多轮设置中运行,使得多轮越狱成为现实的威胁模型和自动化红队测试的重要设置。学习多轮越狱攻击者的一个核心挑战是信用分配:不同轮次对最终结果的贡献不同,但现有学习信号往往过于粗糙,无法识别它们的个体贡献。我们提出了分解信用GRPO(DC-GRPO),这是一个用于多轮越狱学习中组相对策略优化的统一轮次级信用分配框架。DC-GRPO通过结合即时和未来信用为每个轮次分配单独的组相对学习信号,避免了在对话中广播单个轨迹级分数所导致的信用错误分配。我们用静态和动态加权规则实例化这个框架,它们在平衡两种信用来源的方式上有所不同,但共享相同的轮次级结构。在多个受害者大语言模型和基准测试中,动态加权和静态加权变体的平均ASR5@3分数分别达到98.26%和97.88%,显著优于包括SEMA(86.58%)和TROJail(86.23%)在内的现有方法。它们持续强劲的性能表明,核心的经验益处来自轮次级组相对信用分配,而不是特定的加权规则。警告:本文包含有害内容的示例。
英文摘要
Modern large language models (LLMs) operate in interactive multi-turn settings, making multi-turn jailbreaking a realistic threat model and an important setting for automated red teaming. A core challenge in learning multi-turn jailbreak attackers is credit assignment: different turns contribute differently to the final outcome, yet existing learning signals are often too coarse to identify their individual contributions. We propose decomposed credit GRPO (DC-GRPO), a unified turn-level credit assignment framework for Group Relative Policy Optimization in multi-turn jailbreak learning. DC-GRPO assigns a separate group-relative learning signal to each turn by combining immediate and future credit, avoiding the credit misassignment induced by broadcasting a single trajectory-level score across the dialogue. We instantiate this framework with static and dynamic weighting rules that differ in how the two credit sources are balanced while sharing the same turn-level structure. Across multiple victim LLMs and benchmarks, the dynamic- and static-weighted variants achieve average ASR5@3 scores of 98.26% and 97.88%, respectively, substantially outperforming the state-of-the-art methods, including SEMA (86.58%) and TROJail (86.23%). Their consistently strong performance indicates that the central empirical benefit comes from turn-level group-relative credit assignment rather than a particular weighting rule. Warning: This paper contains examples of harmful content.
发表机构
- POSTECH GSAI(浦项科技大学全球人工智能学院)
- Samsung SDS(三星 SDS)
机构由 AI 辅助整理,请以论文原文为准。