GUARD: Role-playing to Generate Natural-language Jailbreakings to Test Guideline Adherence of Large Language Models
Haibo Jin, Ruoxi Chen, Peiyan Zhang, Andy Zhou, Haohan Wang
机构
*
School of Information Sciences University of Illinois at Urbana-Champaign(信息科学学院伊利诺伊大学厄巴纳-香槟分校)
;
Independent Researcher, Starc Institute(Starc研究所独立研究者)
;
Computer Science and Engineering HKUST(HKUST计算机科学与工程学院)
;
Computer Science Lapis Labs University of Illinois Urbana-Champaign(计算机科学Lapis Labs伊利诺伊大学厄巴纳-香槟分校)
;
School of Information Sciences University of Illinois Urbana-Champaign(信息科学学院伊利诺伊大学厄巴纳-香槟分校)
机构
*
University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
University of Pennsylvania(宾夕法尼亚大学)
;
University of California San Diego(加州大学圣地亚哥分校)
;
University of Michigan(密歇根大学)
;
Amazon AGI(亚马逊人工通用智能)
CommentsTo Appear in NeurIPS 2025 Datasets & Benchmarks Track