SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models
SOM方向优于单一方向:语言模型中的多方向拒绝抑制
AI总结 本文提出利用自组织映射(SOM)提取多方向拒绝特征,通过分析有害提示表示与无害提示表示的差异,验证了多方向抑制方法在提升模型安全性和拒绝能力上的有效性。
Comments Accepted at AAAI 2026
Journal ref Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), 2026