Mechanistic Interpretability of Antibody Language Models Using SAEs
使用 SAE 对抗体语言模型的机制可解释性研究
机构 * Department of Statistics, University of Oxford, UK(英国牛津大学统计系) ; Reticular, San Francisco, USA(美国旧金山Reticular公司) ; EECS, MIT, Cambridge MA, USA(美国麻省理工学院电子工程与计算机科学系) ; Leyden Laboratories BV, Leiden, The Netherlands(荷兰莱顿实验室)
专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI、cs.LG
AI总结 本研究采用 TopK 和 Ordered 稀疏自编码器(SAE)对抗体语言模型进行机制可解释性分析,发现 TopK SAE 能揭示有意义的生物学潜在特征但无法保证生成控制,而 Ordered SAE 通过层次结构可靠识别可操控特征但激活模式更复杂。
Comments v3: 15 pages; corrected author list and affiliations in the main text; minor text changes; updated steering results following minor code changes; conclusions and findings remain unchanged; included link to data and code in the Data Availability section