Neural Transparency: Mechanistic Interpretability Interfaces for Anticipating Model Behaviors for Personalized AI
神经透明:用于预测模型行为的机制可解释性接口以实现个性化AI
机构 * MIT Media Lab(麻省理工学院媒体实验室) ; Massachusetts Institute of Technology(麻省理工学院)
专题命中 提示注入 :safety(abstract);分类 cs.AI
AI总结 本研究提出神经透明接口,通过可视化语言模型内部结构帮助用户预测AI行为,提升信任并促进更安全的人机交互。
Comments SK and AB are co-first authors