发表机构
Capital One; Northeastern University(Capital One; 东北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文利用稀疏自编码器分析文本分类语言模型的内部特征,提出探索特征与预测错误关系的技术,并集成到工具SAEfarer中,经专家评估验证其有效性。
AI 中文摘要
随着语言模型(LMs)的日益突出,人们对其透明度产生了兴趣,以便更好地理解其内部行为。最近的可解释性工作集中在使用稀疏自编码器(SAEs)将语言模型中某一层的神经元激活分解为人类可理解的特征,其中每个特征代表模型学到的概念。在本文中,我们分享了使用SAEs分析文本分类语言模型行为的工作。我们提出了探索SAE特征与模型预测和错误之间关系的技术。我们将这些技术整合到SAEfarer中,这是一个用于分析文本分类语言模型所学概念的工具。我们通过五位博士生的专家试点评估来评估SAEfarer。
英文摘要
As language models (LMs) rise in prominence, there is interest in making them more transparent in order to better understand their internal behavior. Recent interpretability work has focused on using sparse autoencoders (SAEs) to break down neuron activations at a given layer in the LM into human-understandable features, where each feature represents a concept that the model has learned. In this paper, we share work on using SAEs to analyze the behavior of text classification LMs. We present techniques for exploring the relationships between the SAE's features and the model's predictions and errors. We integrate these techniques into SAEfarer, a tool for analyzing concepts learned by text classification LMs. We assess SAEfarer in an expert pilot evaluation with five Ph.D. students.
Comments11 pages, 13 figures. Accepted as a short paper at IEEE VIS 2026