Do we really have to filter out random noise in pre-training data for language models?
机构 * School of Electronic and Computer Engineering, Peking University(北京理工大学电子与计算机工程学院) ; University of Electronic Science and Technology of China(电子科技大学) ; Hong Kong University of Science and Technology(香港理工大学) ; Sichuan University(四川大学)
专题命中 预训练与数据 :language model(title);分类 cs.CL