Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation
机构 * Boise State University(博伊州立大学) ; University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校) ; University of Pittsburgh(匹兹堡大学) ; New York University(纽约大学) ; Florida International University(佛罗里达国际大学)
专题命中 安全训练 :safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI
Comments Accepted for EMNLP'25 Findings. TL;DR: We propose a new two-stage intent-based prompt-refinement framework, IntentPrompt, that aims to explore the vulnerability of LLMs' content moderation guardrails by refining prompts into benign-looking declarative forms via intent manipulation for red-teaming purposes