MoH: Multi-Head Attention as Mixture-of-Head Attention
MoH:多头注意力作为专家混合注意力
机构 * School of Electronic and Computer Engineering, Shenzhen Graduate School, Peking University, Shenzhen, China(电子与计算机工程系,深圳研究生院,北京大学,深圳,中国) ; Pengcheng Laboratory, Shenzhen, China(鹏城实验室,深圳,中国) ; School of AI for Science, Shenzhen Graduate School, Peking University, Shenzhen, China(科学人工智能学院,深圳研究生院,北京大学,深圳,中国) ; National University of Singapore, Singapore(新加坡国立大学,新加坡)
AI总结 MoH通过将注意力头视为专家,提升推理效率并优化性能,仅使用部分注意力头即可超越传统多头注意力。
Comments Accepted by ICML 2025, code: https://github.com/SkyworkAI/MoH