arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17523cs.NI

完成路径 credits:面向规模化 Fabric 的多资源控制

Completion-Path Credits: Multi-Resource Control for Scale-Up Fabrics

Fan Yang, Jiaqi Liu, Tao Jiang, Zhan Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对规模化 Fabric 中 credits 难以表征小型操作的问题,提出 SemaCredit 接收端控制器,在保障 HBM 流量性能的同时,显著降低原子操作竞争、响应内爆等场景下的小型操作 P99 延迟,适配不同类型的应用流量。

中文摘要 AI 辅助

规模化 Fabric 用于连接 GPU 和 AI 加速器,承载张量传输以及远程读写、原子操作、通知等操作,这些操作均通过共享的目标侧接收端资源完成。以字节为单位的 credits 可保护链路缓冲区和流式 HBM 流量,但难以准确表征由原子操作(Atomic)执行或响应注入主导的小型操作。本文提出 SemaCredit,这是一种接收端控制器,它会根据目标资源需求向量接纳每个远程内存操作,并在对应 HBM、原子操作或响应阶段完成后返回各组件。在具备多路径队列、8 个 HBM 分区、序列化原子操作引擎和响应引擎的确定性事件模拟器中,SemaCredit 在 HBM 热点流量上与强大的单资源字节基线表现相当,同时在原子操作竞争下将小型操作的 P99 延迟降低 52.4%,在响应内爆下降低 10.2%;在应用驱动的混合流量中,针对 AllReduce 型和远程读取型流量分别实现 57.7% 和 14.5% 的 P99 延迟改善,且在 HBM 主导的 MoE 流量上与字节 credits 表现一致。

英文摘要

Scale-up fabrics connecting GPUs and AI accelerators carry tensor transfers together with remote reads, writes, atomics, and notifications over shared target-side receiver resources. Byte-denominated credits protect link buffers and streaming HBM traffic, but poorly represent small operations dominated by Atomic execution or response injection. This paper presents SemaCredit, a receiver controller that admits each remote-memory operation against a vector of target-resource demands and returns each component when its corresponding HBM, Atomic, or response stage completes. In a deterministic event simulator with multipath queues, eight HBM partitions, a serialized Atomic engine, and a response engine, SemaCredit matches a strong per-resource byte baseline on HBM-hotspot traffic while reducing small-operation P99 latency by 52.4% under Atomic contention and 10.2% under response incast. Application-shaped mixes show 57.7% and 14.5% P99 latency improvements for AllReduce-shaped and remote-read-shaped traffic while matching byte credits on HBM-dominated MoE traffic.

补充信息

↑