Prototype Transformer: Towards Language Model Architectures Interpretable by Design
原型Transformer:迈向可解释设计的语言模型架构
机构 * University of Cambridge(剑桥大学) ; ETH Zurich(苏黎世联邦理工学院)
专题命中 其他推理 :reasoning(abstract);分类 cs.CL、cs.AI、cs.LG
AI总结 提出原型Transformer(ProtoT),一种用线性代价原型模块替代二次代价自注意力的自回归语言模型架构,原型自动捕获可命名概念,提升可解释性并支持行为编辑。
Comments Accepted at ICML 2026. Equal contribution: Yordan Yordanov and Matteo Forasassi. 40 pages, 28 figures, 22 tables