arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2603.08343cs.LGcs.CL

重新思考注意力输出投影:用于高效Transformer的结构化Hadamard变换

Rethinking Attention Output Projection: Structured Hadamard Transforms for Efficient Transformers

Shubham Aggarwal, Lokendra Kumar

首次发表 更新
浏览论文内容

中文总结 AI 辅助

本文提出用固定无参数的Walsh-Hadamard变换替代多头注意力的密集输出投影,减少参数并提升效率,通过实验验证结构化变换在不同复杂度下优于密集投影。

中文摘要 AI 辅助

多头注意力中的密集输出投影随着模型维度平方增长,显著增加参数数量、内存占用和推理成本。本文提出用固定无参数的Walsh-Hadamard变换(WHT)后接对角线仿射变换替代该投影。此方法每块减少约25%的注意力参数,通过正交、保持范数的变换维持全局跨头交互。实验表明,WHT增强模型在验证损失曲线相对于训练FLOPs表现更优,表明训练期间计算利用更高效。关键发现是效率提升如内存占用减少和吞吐量增加随着模型大小、批量大小和序列长度单调增长。本文在prefill和解码阶段评估性能,发现结构化变换在复杂度增加时始终优于密集投影。研究结果表明,用结构化变换替代密集投影可实现更高效的架构,在等效训练预算下获得更低的损失。

英文摘要

The dense output projection in multi head attention scales quadratically with model dimension, contributing significantly to parameter count, memory footprint, and inference cost. We propose replacing this projection with a fixed, parameter free Walsh Hadamard Transform (WHT) followed by a diagonal affine transformation. This approach eliminates approximately 25 percent of attention parameters per block while maintaining global cross-head interaction through an orthogonal, norm-preserving transformation. Our results demonstrate that WHT augmented models exhibit a steeper validation loss curve relative to training FLOPs compared to dense baselines, suggesting superior compute utilization during training. Crucially, we show that efficiency gains including reduced memory footprint and increased throughput grow monotonically with model size, batch size, and sequence length. We evaluate performance across both prefill and decoding stages, finding that the structured transform consistently outperforms dense projections as complexity increases. Our findings indicate that replacing dense projections with structured transforms allows for more compute-efficient architectures that achieve lower loss than dense models at an equivalent training budget.

补充信息

↑