Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges
机构 * Department of Statistics, University of Georgia(统计学系,佐治亚大学) ; School of Computing, University of Georgia(计算学院,佐治亚大学) ; Department of Biostatistics, Boston University(生物统计学系,波士顿大学) ; Department of Epidemiology & Biostatistics, University of Georgia(流行病学与生物统计学系,佐治亚大学) ; School of Computer and Cyber Sciences, Augusta University(计算机与网络科学学院,奥古斯塔大学) ; School of Electrical and Computer Engineering, University of Georgia(电气与计算机工程学院,佐治亚大学) ; Department of Statistics & Data Science, University of Arizona(统计学与数据科学系,亚利桑那大学) ; Kellogg School of Management, Northwestern University(凯洛格管理学院,西北大学) ; Department of Statistics, Harvard University(统计学系,哈佛大学) ; School of Computer Science, Carnegie Mellon University(计算机科学学院,卡内基梅隆大学)
专题命中 指令微调 :large language model(title,abstract);language model(title,abstract);LLM(abstract);preference optimization(abstract)
Comments 119 pages, 10 figures, 7 tables