Safety Pretraining: Toward the Next Generation of Safe AI
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
专题命中 安全训练 :safety(title,abstract);alignment(abstract);分类 cs.LG
机构 * School of Computing, Macquarie University(计算机学院,麦考瑞大学)
专题命中 安全训练 :alignment(abstract);safety(abstract);harmlessness(abstract);分类 cs.CL
Comments EMNLP'25 Main
专题命中 安全训练 :prompt injection(abstract);分类 cs.CL、cs.AI
机构 * Department of Computer Science and Engineering, UNIST(UNIST计算机科学与工程系) ; Graduate School of Artificial Intelligence, POSTECH(POSTECH人工智能研究生院) ; Department of Computer Science and Engineering, POSTECH(POSTECH计算机科学与工程系)
专题命中 安全训练 :safety(abstract);分类 cs.CL、cs.AI
Comments Findings of EMNLP 2025 (32 pages). Code available at https://github.com/seokhyunan/response-tuning
机构 * Fudan University(复旦大学) ; ByteDance Inc(字节跳动公司) ; University of New South Wales(新南威尔士大学)
专题命中 安全训练 :alignment(abstract);分类 cs.CL
Comments Accepted by EMNLP 2025 at the Main Conference
机构 * African Institute for Mathematical Sciences(非洲数学科学研究所)
专题命中 安全训练 :safety(abstract);分类 cs.LG
Comments Submitted to Workshop on AI and ML for Next-Generation Wireless Communications and Networking, NeurIPS 2025
专题命中 安全训练 :safety(abstract)