From Narrow Unlearning to Emergent Misalignment: Causes, Consequences, and Containment in LLMs
机构 * University of Southern California(南加州大学) ; Amazon AGI(亚马逊人工智能实验室)
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG
AI 大模型
大模型对齐、安全、越狱、红队、提示注入和可信评测。
机构 * University of Southern California(南加州大学) ; Amazon AGI(亚马逊人工智能实验室)
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI、cs.LG
机构 * Barclays, Model Risk Management(巴克莱银行,模型风险管理部门) ; Columbia University(哥伦比亚大学)
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.CL
Comments AAAI 2026 Oral. 14 pages (including appendix), 11 figures. Code, data, results, and additional resources are available at: https://model-editing.github.io
专题命中 AI治理与伦理 :alignment(abstract);safety(abstract);分类 cs.AI
Comments 12 pages, 4 figures, 1 table, includes Supplementary Materials, simulation code on GitHub (https://github.com/AerisSpace/SecondLawIntelligence )
机构 * Department of Computer Science and Engineering, BRAC University(计算机科学与工程系,BRAC大学)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL、cs.LG
机构 * Cornell Tech(康奈尔科技) ; Microsoft Research(微软研究院)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI、cs.LG
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.CL
Comments To be published in the Proceedings of the 40th Annual AAAI Conference on Artificial Intelligence (AAAI 2026 Special Track on AI for Social Impact )
机构 * Microsoft Corporation(微软公司)
专题命中 AI治理与伦理 :alignment(abstract);分类 cs.AI