arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RED-PIM:使用内存内处理减少Transformer的数据移动

RED-PIM: Reducing Data Movement for Transformers using Processing-in-Memory

Zahra Yousefijamarani, Alaa Alameldeen

arXiv 2607.21731首次发表:更新:

AI 中文总结

针对Transformer注意力操作数据移动量大影响效率问题,提出RED-PIM算法架构协同设计,通过减少银行间数据移动、缩小中间注意力矩阵,重组操作、本地计算和优化传输策略,降低成本与流量,提升推理性能。

AI 中文摘要

Transformer广泛应用于自然语言处理、计算机视觉、网络搜索和DNA序列分析等多个领域。提高其性能至关重要。注意力操作期间处理单元与内存间大量的数据移动限制了效率。内存内处理(PIM)可缓解此问题,但此前基于PIM的Transformer实现存在银行间通信成本高、内存银行容量有限难以扩展等问题。本文提出RED-PIM,通过将银行间数据移动从O(N^2)降至O(N),将中间注意力矩阵从N x N缩小到d x d来减少注意力延迟。通过重组矩阵操作、本地计算和优化数据传输策略,RED-PIM显著降低计算成本和互连流量。与基线PIM实现相比,RED-PIM推理时间减少16.05%至99.99%(几何平均值66.42%),在实际数据集上,长文档性能提高99.60%,短文档提高13.44%,且保持或提高了准确性,证明了RED-PIM对可扩展高效Transformer推理的有效性。

英文摘要

Transformers are widely used across many domains, including natural language processing, computer vision, web search, and DNA sequence analysis. Given their broad applicability, improving the performance of transformer models is critical. However, the high volume of data movement between processing units and memory during attention operations significantly limits their efficiency. Processing-In-Memory (PIM) mitigates this issue by performing computations directly inside memory. While prior work has proposed PIM-based transformer implementations, they suffer from costly inter-bank communication, and struggle to scale due to the limited capacity of memory banks. As a result, attention-related data must be split across banks, diminishing the potential benefits of PIM. In this work, we propose RED-PIM, an algorithm-architecture co-design that reduces attention latency by minimizing inter-bank data movement from O(N^2) to O(N) and shrinking intermediate attention matrices from N x N to d x d. By reorganizing matrix operations, performing computations locally, and employing an optimized data transfer strategy, RED-PIM significantly reduces computation cost and interconnect traffic. Compared to baseline PIM implementation, RED-PIM achieves inference time reductions ranging from 16.05% to 99.99% (geometric mean of 66.42%), with the largest gains on longer sequences. On real-world datasets, RED-PIM improves performance by 99.60% for long documents and 13.44% for shorter ones, while maintaining or improving accuracy. These results demonstrate RED-PIM's effectiveness for scalable and efficient transformer inference.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑