发表机构
IMEC(imec)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过封装级热模型量化了3D立体DRAM-on-GPU集成的热设计包络,发现堆叠高度是峰值温度主因,冷却腔电导率等参数调节可行性,并定义了带宽、容量与热约束下的协同设计空间。
AI 中文摘要
用于AI和HPC工作负载的GPU扩展日益受到2.5D HBM-GPU和直接堆叠的3D HBM-on-GPU集成在容量、带宽和热限制方面的制约。本研究建立了3D立体DRAM-on-GPU集成的热设计包络,其中垂直取向的DRAM芯片和交错冷却腔重塑了GPU上方的热流和存储器接口。使用锚定于一致HBM-on-GPU基线并由实际光罩尺度非均匀GPU功率图驱动的封装级热模型,我们量化了控制热可行性的关键参数。堆叠高度是峰值温度的主要限制因素,而冷却腔电导率改变了可行区域,模具插入和堆叠取向进一步调节热行为。分布式存储控制器和片上网络层仅引入适度的热惩罚。尽管芯片级并行性增加了带宽,但一旦执行变为计算受限,模拟训练时间的减少趋于饱和。这些结果为3D立体DRAM-on-GPU集成在带宽、容量和热约束下定义了一个有界的协同设计空间。
英文摘要
The scaling of GPUs for AI and HPC workloads is increasingly constrained by the capacity, bandwidth, and thermal limits of both 2.5D HBM-GPU and direct-stacked 3D HBM-on-GPU integration. This work establishes the thermal design envelope for 3D volumetric DRAM-on-GPU integration, in which vertically oriented DRAM dies and interleaved cooling cavities reshape heat flow and memory interfacing above the GPU. Using a package-level thermal model anchored to a consistent HBM-on-GPU baseline and driven by a realistic reticle-scale non-uniform GPU power map, we quantify the key parameters governing thermal feasibility. Stack height is the dominant limiter of peak temperature, while cooling-cavity conductivity shifts the feasible region, and mold insertion and stack orientation further modulate thermal behavior. A distributed memory-controller and network-on-chip tier introduces only a moderate thermal penalty. Although die-level parallelism increases bandwidth, the reduction in simulated training time saturates once execution becomes compute-bound. These results define a bounded co-design space across bandwidth, capacity, and thermal constraints for 3D volumetric DRAM-on-GPU integration.
CommentsPresented at the 52nd IEEE European Solid-State Electronics Research Conference (ESSERC 2026), Palma de Mallorca, Spain, September 7-10, 2026. To appear in the conference proceedings