arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PsychJail:通过对大语言模型策略的多轮说服探索心理越狱

PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies

Zeyu Feng, Qingyu Wu, Yuzhe Luo, Hua Cheng

arXiv 2608.23028首次发表:更新:

发表机构

The Defense Innovation Institute, Academy of Military Sciences(军事科学院国防创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出PsychJail框架,利用社会心理学技术构建攻击策略,通过轨迹级强化学习优化,在四个对齐LLM上实现87.3%平均攻击成功率,揭示了不同模型的心理画像指纹,为LLM安全红队提供了新方向。

AI 中文摘要

大语言模型(LLM)正越来越多地被部署在教育、医疗、政策咨询及其他交互场景中,用户将其作为持续的社交对话者而非单次查询引擎使用。这一转变使得越狱成为日益严峻的安全威胁,但多数研究聚焦于单轮提示优化或迭代攻击优化,忽视了基于心理学的多轮漏洞。本文提出PsychJail,一种基于心理学的红队框架,通过符合理论的多轮说服来攻击对齐后的LLM。PsychJail将成熟的社会心理学说服技术映射为策略条件化的攻击策略,将每个攻击者行为分解为意义改变分析、策略选择和受害者可见消息,实现了说服知识模型(PKM)。该策略通过轨迹级强化学习优化,使用PKM门控奖励,仅当每轮都包含结构完整的意义改变分析时才计入早期越狱成功。在四个对齐后的受害者模型上,PsychJail达到了最高的平均攻击成功率(87.3%),且在每个模型上都优于强大的单轮和多轮基线。我们还测量了每个受害者被突破时的行为敏感性,揭示了四种不同的模型级指纹,可识别哪些说服手段影响每个模型及影响范围,这些指纹有助于解释跨模型迁移不对称性。我们将其解释为四种候选心理画像——理性主义者、可信度驱动型、叙事单文化型和广泛可说服型,同时将此解释视为需要未来验证的猜想。我们的研究结果确立了心理越狱作为日益交互的LLM的独特红队前沿。

英文摘要

Large language models (LLMs) are increasingly deployed in education, healthcare, policy advising, and other interactive settings, where users engage them as sustained social interlocutors rather than one-shot query engines. This shift makes jailbreaks a growing safety threat, yet most research emphasizes single-turn prompt optimization or iterative attack refinement, leaving psychologically grounded multi-turn vulnerabilities underexplored. We present PsychJail, a psychology-guided framework for red teaming aligned LLMs through theory-grounded, multi-turn persuasion. PsychJail maps established social-psychological persuasion techniques into a tactic-conditioned attack policy. It factorizes each attacker action into a Change-of-Meaning analysis, tactic selection, and victim-visible message, operationalizing the Persuasion Knowledge Model (PKM). The policy is refined with trajectory-level reinforcement learning using a PKM-gated reward that credits early jailbreak success only when every turn contains a well-formed Change-of-Meaning analysis. Across four aligned victim models, PsychJail achieves the highest average attack success rate (87.3%) and outperforms strong single-turn and multi-turn baselines on every model. We also measure susceptibility at the action that breaks each victim, revealing four distinct model-level fingerprints that identify which persuasion levers affect each model and how broadly. These fingerprints help explain cross-model transfer asymmetry. We interpret them as four candidate psychological profiles-rationalist, credibility-driven, narrative-monoculture, and broadly persuadable-while treating this interpretation as a conjecture requiring future validation. Our findings establish psychological jailbreaks as a distinct red-teaming frontier for increasingly interactive LLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑