机构
*
School of Mechanical and Electrical Engineering, University of Electronic Science and Technology of China(电子科技大学机械与电子工程学院)
;
Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
;
Department of Pathology, Sichuan Clinical Research Center for Cancer, Sichuan Cancer Hospital & Institute, Affiliated Cancer Hospital of University of Electronic Science and Technology of China(pathology department, 四川省癌症临床研究中心, 四川省肿瘤医院及研究所, 电子科技大学附属肿瘤医院)
;
Department of Radiation Oncology, Sichuan Cancer Hospital and Institute, University of Electronic Science and Technology of China(放射肿瘤科, 四川省肿瘤医院及研究所, 电子科技大学)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
Abhay Sheshadri, Aidan Ewart, Phillip Guo, Aengus Lynch, Cindy Wu, Vivek Hebbar, Henry Sleight, Asa Cooper Stickland, Ethan Perez, Dylan Hadfield-Menell, Stephen Casper
机构
*
Georgia Institute of Technology(佐治亚理工学院)
;
University of Bristol(布里斯托大学)
;
University of Maryland(马里兰大学)
;
University College London(伦敦大学学院)
;
MATS
;
Astra
;
MIT CSAIL(麻省理工学院计算机科学与人工智能实验室)
专题命中
知识编辑与模型理解
:LLM(abstract,comments);large language model(abstract);language model(abstract);分类 cs.CL、cs.AI、cs.LG
Can sparse autoencoders make sense of gene expression latent variable models?
Viktoria Schuster
机构
*
Eric and Wendy Schmidt Center, Broad Institute of MIT and Harvard(埃里克和文迪斯中心,哈佛-麻省理工Broad研究所)
;
Department of Computer Science, University of Copenhagen(哥本哈根大学计算机科学系)
专题命中
知识编辑与模型理解
:large language model(abstract);language model(abstract);分类 cs.LG