Continual Pretraining on Encrypted Synthetic Data for Privacy-Preserving LLMs
在加密合成数据上进行持续预训练以实现隐私保护的大语言模型
机构 * The PII shown in the figure is synthetic and not real(合成的PII) ; International Digital Economy Academy(国际数字经济学院) ; The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ; The Hong Kong University of Science and Technology(香港科学与技术大学) ; DataArc Tech Ltd(DataArc科技有限公司)
专题命中 预训练与数据 :pretraining(title,abstract);large language model(abstract);language model(abstract);分类 cs.CL
AI总结 本文提出了一种基于实体的加密合成数据预训练框架,通过加密保护PII,实现隐私保护的大语言模型预训练,同时保持模型的指令遵循能力。