arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03061cs.AR

分而治之:MCM GPU中的可扩展性能与能耗

Divide and conquer: Scalable performance and energy in MCM GPUs

Mario Ibáñez Bolado, Borja Pérez Pavón, Jose Luis Bosque Orero, Julio Ramón Beivide

首次发表
浏览论文内容

中文总结 AI 辅助

本文研究MCM GPU中跨小芯片分布计算与存储的可扩展性,证明16小芯片环形拓扑配256个SM相比同算力最先进架构性能提升2.40倍、能耗降低4.45倍,凸显解聚作为关键架构因素的价值。

中文摘要 AI 辅助

多芯片模块(MCM)GPU通过在公共封装上集成多个小芯片,为超越单片设计的计算能力扩展提供了一条有前景的路径。然而,解聚对性能可扩展性和能耗的影响仍未得到充分探索。设计空间在每小芯片SM数、小芯片数量和互连网络等维度上迅速增长。小芯片间互连网络尤为关键,因为它决定了额外的计算资源能否转化为性能提升。这种有限的理解使得工业界和研究界在MCM GPU扩展的性能与能耗权衡方面缺乏明确指导。在本工作中,我们研究了将计算和存储能力分布到多个小芯片上是否比集中资源提供更可扩展的替代方案。我们量化了它们对性能、能耗和效率的影响,并考察了小芯片间拓扑如何在不同系统规模下影响可扩展性。我们的结果表明,具有256个SM的16小芯片环形拓扑配置相比具有相同计算能力的最先进MCM架构,实现了显著的$2.40\times$性能提升,同时能耗降低了$4.45\times$。这些显著收益证明了解聚是一阶架构因素,对于释放下一代GPU的性能和能效潜力至关重要。

英文摘要

Multi-chip-module (MCM) GPUs offer a promising path to scale compute capability beyond monolithic designs by integrating multiple chiplets on a common package. However, the impact of disaggregation on performance scalability and energy consumption remains underexplored. The design space grows rapidly across dimensions such as SMs per chiplet, chiplet count, and interconnection network. The inter-chiplet network is particularly critical, as it determines whether additional compute resources translate into performance gains. This limited understanding leaves industry and research without clear guidance on the performance and energy trade-offs of MCM GPU scaling. In this work, we investigate whether distributing compute and memory capability across multiple chiplets offers a more scalable alternative to concentrating resources. We quantify their effects on performance, energy, and efficiency and examine how inter-chiplet topology influences scalability at different system sizes. Our results demonstrate that a 16 chiplet Torus configuration with 256 SMs delivers a remarkable $2.40\times$ performance improvement over a state-of-the-art MCM architecture with the same compute capability, while simultaneously reducing energy consumption by $4.45\times$. These substantial gains provide evidence that disaggregation is a first-order architectural factor and will be critical to unlocking the performance and energy-efficiency potential of next-generation GPUs.

发表机构

  • Universidad de Cantabria(坎塔布里亚大学)

机构由 AI 辅助整理,请以论文原文为准。

↑