Antidote: Post-fine-tuning Safety Alignment for Large Language Models against Harmful Fine-tuning
机构 * Georgia Institute of Technology(佐治亚理工学院) ; Dolby Laboratories(杜比实验室)
专题命中 安全评测 :alignment(title,abstract);safety(title,abstract);分类 cs.AI
Comments Rejected by AAAI25-AIA. Accepted by ICML25. Authors are thankful to the anonymous reviewers from both AAAI25-AIA and ICML25