完成路径 credits:面向规模化 Fabric 的多资源控制
Completion-Path Credits: Multi-Resource Control for Scale-Up Fabrics
浏览论文内容
中文总结 AI 辅助
针对规模化 Fabric 中 credits 难以表征小型操作的问题,提出 SemaCredit 接收端控制器,在保障 HBM 流量性能的同时,显著降低原子操作竞争、响应内爆等场景下的小型操作 P99 延迟,适配不同类型的应用流量。
中文摘要 AI 辅助
规模化 Fabric 用于连接 GPU 和 AI 加速器,承载张量传输以及远程读写、原子操作、通知等操作,这些操作均通过共享的目标侧接收端资源完成。以字节为单位的 credits 可保护链路缓冲区和流式 HBM 流量,但难以准确表征由原子操作(Atomic)执行或响应注入主导的小型操作。本文提出 SemaCredit,这是一种接收端控制器,它会根据目标资源需求向量接纳每个远程内存操作,并在对应 HBM、原子操作或响应阶段完成后返回各组件。在具备多路径队列、8 个 HBM 分区、序列化原子操作引擎和响应引擎的确定性事件模拟器中,SemaCredit 在 HBM 热点流量上与强大的单资源字节基线表现相当,同时在原子操作竞争下将小型操作的 P99 延迟降低 52.4%,在响应内爆下降低 10.2%;在应用驱动的混合流量中,针对 AllReduce 型和远程读取型流量分别实现 57.7% 和 14.5% 的 P99 延迟改善,且在 HBM 主导的 MoE 流量上与字节 credits 表现一致。
英文摘要
Scale-up fabrics connecting GPUs and AI accelerators carry tensor transfers together with remote reads, writes, atomics, and notifications over shared target-side receiver resources. Byte-denominated credits protect link buffers and streaming HBM traffic, but poorly represent small operations dominated by Atomic execution or response injection. This paper presents SemaCredit, a receiver controller that admits each remote-memory operation against a vector of target-resource demands and returns each component when its corresponding HBM, Atomic, or response stage completes. In a deterministic event simulator with multipath queues, eight HBM partitions, a serialized Atomic engine, and a response engine, SemaCredit matches a strong per-resource byte baseline on HBM-hotspot traffic while reducing small-operation P99 latency by 52.4% under Atomic contention and 10.2% under response incast. Application-shaped mixes show 57.7% and 14.5% P99 latency improvements for AllReduce-shaped and remote-read-shaped traffic while matching byte credits on HBM-dominated MoE traffic.