When the Coffee Feature Activates on Coffins: An Analysis of Feature Extraction and Steering for Mechanistic Interpretability
当咖啡特征在棺材上激活:对特征提取和转向用于机制可解释性的分析
机构 * Department of Philosophy of Nature and Technology(自然哲学与技术系) ; Munich School of Philosophy(慕尼黑哲学学院) ; Division of the Humanities and Social Sciences(人文与社会科学系) ; California Institute of Technology(加州理工学院)
AI总结 本文分析了通过稀疏自编码器提取特征和控制模型输出的方法,指出其在机制可解释性中的局限性和可靠性问题,强调需转向更可靠的预测与控制。
Comments 33 pages (65 with appendix), 1 figure