SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment
机构 * Department of Artificial Intelligence, Yonsei University(人工智能系,延世大学) ; Department of Computer Science and Engineering, Yonsei University(计算机科学与工程系,延世大学)
专题命中 安全训练 :alignment(title,abstract);safety(abstract);jailbreak(abstract);分类 cs.CL、cs.AI
Comments Accepted at NeurIPS 2025. Code and models are available at https://ai-isl.github.io/safepath