From Logits to Latents: Contrastive Representation Shaping for LLM Unlearning
从 logits 到 latents:用于大语言模型遗忘的对比表示塑造
机构 * Purdue University(普渡大学)
专题命中 隐私与版权 :alignment(abstract);分类 cs.LG
AI总结 CLReg 通过对比表示正则化减少大语言模型中遗忘与保留知识的纠缠,提升遗忘效果并降低隐私风险。
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
从 logits 到 latents:用于大语言模型遗忘的对比表示塑造
机构 * Purdue University(普渡大学)
专题命中 隐私与版权 :alignment(abstract);分类 cs.LG
AI总结 CLReg 通过对比表示正则化减少大语言模型中遗忘与保留知识的纠缠,提升遗忘效果并降低隐私风险。