arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HBFSim:真实GPU执行下高带宽闪存的快速忠实模拟

HBFSim: Fast and Faithful Simulation of High-Bandwidth Flash Under Real GPU Execution

Yanpeng Hu, Yiwei Yang, Yuanwu Zhu, Yusheng Zheng, Andi Quinn, Wei Zhang

arXiv 2609.09800首次发表:更新:

发表机构

ShanghaiTech University; UC Santa Cruz; University of Science and Technology of China; University of Connecticut(上海科技大学; 加州大学圣克鲁兹分校; 中国科学技术大学; 康涅狄格大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

HBFSim首个在真实GPU执行下模拟高带宽闪存时序、容量和热效应的平台,通过重写PTX实现,与设备精确匹配,比详细路径快20.8倍。

AI 中文摘要

服务大型语言模型(LLM)受限于内存容量。高带宽闪存(HBF)将NAND闪存堆叠在加速器封装内,位于高带宽内存(HBM)下一层;该规范于2026年8月3日发布,首批推理设备预计在2027年初提供样品。关于容量和数据放置的决策不能等到硅片问世。现有方法均无法解决这些决策:存储模拟器重放记录的访问序列,但从不执行工作负载;GPU模拟器不运行真实计算内核;周期精确模拟器无法完成一次LLM推理运行。我们提出HBFSim,这是第一个在真实推理工作负载于真实GPU上执行时,应用HBF时序、容量和热效应的评估平台。HBFSim重写PTX(NVIDIA编译器发出的中间代码),并门控内核启动;发布与消耗分离,因此真实硬件提供隐藏访问的计算。时序来自真实设备的测量而非参数表,结温同时设定HBF持续速率和强制刷新写入的保留期限。HBFSim在所有六个校准断点处与测量设备完全匹配,零不安全启动,未修改的vLLM 0.15.1服务Qwen3-30B-A3B返回与未插桩基线相同的令牌标识符。设备快速路径在2秒内服务相同的Qwen3-30B-A3B案例,而详细参考路径需44秒,快20.8倍。在HBF部件样品问世之前,HBFSim让设计者能在真实工作负载下衡量容量或放置决策,而非假设。

英文摘要

High-Bandwidth Flash (HBF) places high-capacity NAND beside HBM to relieve the memory-capacity bottleneck of LLM inference, yet its system-level behavior cannot be evaluated before hardware becomes available. Cycle-level GPU simulators are too slow for production-scale models. Trace replay has a further shortcoming: it cannot capture the allocation, migration, and execution changes induced by different HBM-HBF configurations. Our key insight is that HBF need not be evaluated by simulating the GPU: only the program-visible effects of HBF need to be modeled. And only a real LLM workload running on real hardware can answer the arguments about HBF. Hence the modeled service has to be injected into that running program, and the injection must not destroy the GPU concurrency that would hide the original I/O latency. We present HBFSim, an open-source HBF simulator that executes LLM workloads on a real GPU while modeling HBF timing, thermal, and other behaviors online. HBFSim rewrites the PTX of the workload's kernels and routes accesses inside a registered address range into the HBF simulator. It supports asynchronous TMA transfers and capacities beyond physical GPU memory. HBFSim leaves the model's run unaffected across ordinary-memory, TMA, and capacity-mode tests. The delay it injects matches the delay requested to within 0.152%. We also design a coupled thermal module that puts HBF, HBM, and the GPU in one advanced package, which is important for answering how severe the hot throttling problem becomes after HBF runs for a long time. Experiments with Qwen3-30B show how package heating, HBM-HBF allocation, and shared MoE demand jointly constrain the design space of future HBF accelerators. The source code of HBFSim is available at https://github.com/SlugLab/hbfsim/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑