Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
Istabrak Abbes, Gopeshh Subbaraj, Matthew Riemer, Nizar Islah, Benjamin Therien, Tsuguchika Tabaru, Hiroaki Kingetsu, Sarath Chandar, Irina Rish
机构
*
Université de Montréal(蒙特利尔大学)
;
Mila – Quebec AI Institute(魁北克人工智能研究院)
;
Chandar Research Lab(Chandar研究实验室)
;
IBM Research(IBM研究院)
;
Fujitsu Research(富士通研究院)
;
Polytechnique Montréal(蒙特利尔理工学院)
机构
*
School of Medicine, The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)医学院)
;
Shenyang Institute of Automation, Chinese Academy of Sciences(中国科学院沈阳自动化研究所)
;
Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
;
Wave and Machine Intelligence Department, Technology Innovation Institute(技术创新研究院波浪与机器智能部门)
;
University of the Chinese Academy of Sciences(中国科学院大学)
;
McGill University(麦吉尔大学)