Can Editing LLMs Inject Harm?
能否通过编辑LLM注入危害?
机构 * Northwestern University(西北大学)
AI总结 本文提出编辑攻击作为LLM安全威胁,揭示了通过编辑注入虚假信息和偏见的风险及高隐蔽性。
Comments Accepted to Proceedings of AAAI 2026. The first two authors contributed equally. 7 pages for main paper, 31 pages including appendix. The code, results, dataset for this paper and more resources are on the project website: https://llm-editing.github.io