arXivDaily arXiv每日学术速递 周一至周五更新

作者

Dawn Song

AI / Security

2026-01-15 至 2026-01-15 共收录 1
2407.20224 2026-01-15 cs.CL

Can Editing LLMs Inject Harm?

能否通过编辑LLM注入危害?

Canyu Chen, Baixiang Huang, Zekun Li, Zhaorun Chen, Shiyang Lai, Xiongxiao Xu, Jia-Chen Gu, Jindong Gu, Huaxiu Yao, Chaowei Xiao, Xifeng Yan, William Yang Wang, Philip Torr, Dawn Song, Kai Shu

机构 * Northwestern University(西北大学)

AI总结 本文提出编辑攻击作为LLM安全威胁,揭示了通过编辑注入虚假信息和偏见的风险及高隐蔽性。

Comments Accepted to Proceedings of AAAI 2026. The first two authors contributed equally. 7 pages for main paper, 31 pages including appendix. The code, results, dataset for this paper and more resources are on the project website: https://llm-editing.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏