Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron
通过单个神经元的模型自反进行轻量级对齐以提升大语言模型的安全性
机构 * Beijing Institute of AI Safety and Governance(北京人工智能安全与治理研究院) ; Beijing Key Laboratory of Safe AI and Superalignment(北京安全人工智能与超对齐重点实验室) ; Ant Group Co., Ltd.(蚂蚁集团有限公司) ; BrainCog Lab., CASIA(CASIA脑认知实验室) ; Zhongguancun Academy(中关村学院)
专题命中 安全训练 :alignment(title,abstract);safety(title,abstract);分类 cs.AI、cs.LG
AI总结 通过单个神经元的模型自反实现轻量级对齐,提升大语言模型的安全性和实用性
Comments 21 pages, 3 figures