AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models
机构 * The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; State University of New York at Buffalo(纽约州立大学布法罗分校) ; Huawei International, Singapore(新加坡华为国际)
专题命中 安全训练 :alignment(title,abstract);safety(title,abstract);jailbreak(abstract);分类 cs.CL、cs.AI