Large Language Models Encode Semantics and Alignment in Linearly Separable Representations
大语言模型在线性可分的表示中编码语义和对齐
机构 * Yale University(耶鲁大学) ; Foundation AI – Cisco Systems Inc(Foundation AI – 卡西欧系统公司)
专题命中 代码与定理证明 :reasoning(abstract);分类 cs.CL、cs.LG
AI总结 本研究发现大语言模型通过线性可分的表示编码语义和对齐,提出基于潜在空间的MLP探测器有效提升安全防护。
Comments IJCNLP and the Asian Chapter of ACL